TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

December 2024

50 篇论文 · 按点赞排序

33

AniDoc: Animation Creation Made Easier

Yihao Meng, Hao Ouyang, Hanlin Wang +6 authors

AniDoc uses video diffusion models to automate colorization and in-betweening in 2D animation, improving efficiency by leveraging correspondence matching.

58video diffusion modelscorrespondence matchingHF ↗arXiv ↗
35

Parallelized Autoregressive Visual Generation

Yuqing Wang, Shuhuai Ren, Zhijie Lin +6 authors

A parallel generation strategy for autoregressive models improves inference speed without significantly compromising quality in visual generation tasks.

52autoregressive modelsparallelized autoregressiveHF ↗arXiv ↗
36

How to Synthesize Text Data without Model Collapse?

Xuekai Zhu, Daixuan Cheng, Hengli Li +7 authors

The use of synthetic data in language model training leads to model collapse, which is mitigated by token-level editing of human-produced data to create semi-synthetic data.

52synthetic datamodel collapseHF ↗arXiv ↗
39

Multimodal Latent Language Modeling with Next-Token Diffusion

Yutao Sun, Hangbo Bao, Wenhui Wang +5 authors

LatentLM integrates continuous and discrete data using causal Transformers, VAEs, and next-token diffusion, achieving superior performance in multimodal tasks like image generation, large language model integration, and text-to-speech synthesis.

50LatentLMcausal TransformersHF ↗arXiv ↗
40

Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition

Zhisheng Zhong, Chengyao Wang, Yuqi Liu +12 authors

As Multi-modal Large Language Models (MLLMs) evolve, expanding beyond single-domain capabilities is essential to meet the demands for more versatile and efficient AI. However, previous omni-models have insufficiently explored speech, neglecting its integration with multi-modality. We introduce Lyra, an efficient MLLM that enhances multimodal abilities, including advanced long-speech comprehension, sound understanding, cross-modality efficiency, and seamless speech interaction. To achieve efficiency and speech-centric capabilities, Lyra employs three strategies: (1) leveraging existing open-source large models and a proposed multi-modality LoRA to reduce training costs and data requirements; (2) using a latent multi-modality regularizer and extractor to strengthen the relationship between speech and other modalities, thereby enhancing model performance; and (3) constructing a high-quality, extensive dataset that includes 1.5M multi-modal (language, vision, audio) data samples and 12K long speech samples, enabling Lyra to handle complex long speech inputs and achieve more robust omni-cognition. Compared to other omni-methods, Lyra achieves state-of-the-art performance on various vision-language, vision-speech, and speech-language benchmarks, while also using fewer computational resources and less training data.

48Multi-modal Large Language Models (MLLMs)omni-modelsHF ↗arXiv ↗
42

Evaluating and Aligning CodeLLMs on Human Preference

Jian Yang, Jiaxi Yang, Ke Jin +7 authors

A human-curated benchmark (CodeArena) and a large synthetic instruction corpus (SynCode-Instruct) are introduced to evaluate code LLMs based on human preference alignment, revealing performance differences between open-source and proprietary models.

48code large language modelscode generationHF ↗arXiv ↗
46

GRAPE: Generalizing Robot Policy via Preference Alignment

Zijian Zhang, Kaiyuan Zheng, Zhaorun Chen +6 authors

GRAPE improves vision-language-action models' performance by aligning policies via preference modeling, enhancing generalizability and allowing customization of objectives like safety and efficiency.

47vision-language-action modelsGRAPEHF ↗arXiv ↗
47

Token-Budget-Aware LLM Reasoning

Tingxu Han, Chunrong Fang, Shiyu Zhao +3 authors

A token-budget-aware framework dynamically estimates and allocates token budgets for LLM reasoning, reducing costs with minimal performance loss.

46Chain-of-ThoughtCoT reasoningHF ↗arXiv ↗
50

MALT: Improving Reasoning with Multi-Agent LLM Training

Sumeet Ramesh Motwani, Chandler Smith, Rocktim Jyoti Das +6 authors

Multi-agent LLM training improves performance on reasoning tasks by assigning specialized roles and utilizing joint outcome-based rewards to enhance collaboration among models.

46sequential multi-agent setupheterogeneous LLMsHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号