TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jun 5 – Jun 11, 2023
本周最热169

Simple and Controllable Music Generation

Jade Copet, Felix Kreuk, Itai Gat +5 authors

MusicGen, a single-stage transformer language model, generates high-quality music conditioned on text or melodic features using efficient token interleaving patterns, outperforming existing models on text-to-music benchmarks.

Language Modeltransformer LMtoken interleaving patternsHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

Orca: Progressive Learning from Complex Explanation Traces of GPT-4

Subhabrata Mukherjee, Arindam Mitra, Ganesh Jawahar +3 authors

Orca, a 13-billion parameter model, enhances small models by imitating the reasoning process from large foundation models using rich signals and diverse data, surpassing existing models in complex reasoning benchmarks.

51imitation learninglarge foundation modelsHF ↗arXiv ↗
05

MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Bo Li, Yuanhan Zhang, Liangyu Chen +5 authors

MIMIC-IT, a multimodal instruction-response dataset, enhances VLMs' zero-shot performance on vision-language tasks through extensive and diverse data, leading to improved perception, reasoning, and planning.

12MultI-Modal In-Context Instruction TuningMIMIC-ITHF ↗arXiv ↗
06

Recognize Anything: A Strong Image Tagging Model

Youcai Zhang, Xinyu Huang, Jinyu Ma +9 authors

The Recognize Anything Model (RAM) is a powerful image tagging model trained using large-scale image-text pairs, achieving superior zero-shot performance compared to supervised models and Google API.

12foundation modelimage taggingHF ↗arXiv ↗
08

Tracking Everything Everywhere All at Once

Qianqian Wang, Yen-Yu Chang, Ruojin Cai +4 authors

OmniMotion, a globally consistent motion representation using a quasi-3D canonical volume, outperforms state-of-the-art methods in motion estimation by modeling camera and object motion through occlusions.

11optical flowparticle video trackingHF ↗arXiv ↗
10

Segment Anything in High Quality

Lei Ke, Mingqiao Ye, Martin Danelljan +4 authors

HQ-SAM enhances the segmentation capabilities of SAM by introducing a learnable output token that leverages additional Vit features, improving mask quality with minimal computational overhead.

10Segment Anything ModelSAMHF ↗arXiv ↗
15

Matting Anything

Jiachen Li, Jitesh Jain, Humphrey Shi

The Matting Anything Model (MAM) is a single, efficient image matting framework that leverages SAM and a lightweight M2M module for flexible user guidance and performs comparably to specialized models with fewer parameters.

6Matting Anything ModelMAMHF ↗arXiv ↗
17

Emergent Correspondence from Image Diffusion

Luming Tang, Menglin Jia, Qianqian Wang +2 authors

Image diffusion models extract implicit features (DIFT) that outperform supervised and weakly-supervised methods in establishing semantic, geometric, and temporal correspondences between images.

6diffusion modelsDIffusion FeaTures (DIFT)HF ↗arXiv ↗
18

Deductive Verification of Chain-of-Thought Reasoning

Zhan Ling, Yunhao Fang, Xuanlin Li +4 authors

A natural language-based format decomposes reasoning verification into subprocesses, enhancing the rigor and trustworthiness of language model reasoning steps and improving accuracy on complex tasks.

6Chain-of-Thoughtdeductive reasoningHF ↗arXiv ↗
28

MotionDiffuser: Controllable Multi-Agent Motion Prediction using Diffusion

Chiyu Max Jiang, Andre Cornman, Cheolho Park +3 authors

MotionDiffuser, a diffusion model for multi-agent trajectory prediction, achieves state-of-the-art results by learning a multimodal distribution with a simple predictor and permutation-invariant joint learning, using PCA for trajectory compression and differentiable cost functions for constrained sampling.

4diffusion based representationmultimodal distributionHF ↗arXiv ↗
29

HeadSculpt: Crafting 3D Head Avatars with Text

Xiao Han, Yukang Cao, Kai Han +5 authors

HeadSculpt introduces a coarse-to-fine pipeline for generating and editing 3D head avatars with high fidelity and identity preservation using a diffusion model enhanced with 3D awareness and a novel editing score distillation strategy.

4text-guided 3D generative methodstext-to-image diffusion modelHF ↗arXiv ↗
30

PolyVoice: Language Models for Speech to Speech Translation

Qianqian Dong, Zhiying Huang, Chen Xu +14 authors

PolyVoice, a language model-based framework for speech-to-speech translation, uses discretized speech units and VALL-E X for high-quality voice and style preservation in translation.

4language model-based frameworkspeech-to-speech translationHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号