TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

September 2024
本月最热159

Qwen2.5-Coder Technical Report

Binyuan Hui, Jian Yang, Zeyu Cui +14 authors

Qwen2.5-Coder series demonstrates state-of-the-art code generation, completion, reasoning, and repair capabilities using the Qwen2.5 architecture with over 5.5 trillion tokens of training data.

Qwen2.5-CoderQwen2.5-Coder-1.5BQwen2.5-Coder-7BQwen2.5 architectureHF ↗arXiv ↗

50 篇论文 · 按点赞排序

04

OmniGen: Unified Image Generation

Shitao Xiao, Yueze Wang, Junjie Zhou +6 authors

OmniGen is a unified diffusion model for image generation that supports diverse tasks without additional modules, emphasizing simplicity, knowledge transfer, and reasoning capabilities.

115diffusion modelOmniGenHF ↗arXiv ↗
05

Emu3: Next-Token Prediction is All You Need

Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo +22 authors

Emu3, a transformer-based multimodal model trained exclusively with next-token prediction, outperforms existing diffusion and compositional models in generation and perception tasks.

99next-token predictionmultimodal modelsHF ↗arXiv ↗
10

OLMoE: Open Mixture-of-Experts Language Models

Niklas Muennighoff, Luca Soldaini, Dirk Groeneveld +21 authors

A sparse Mixture-of-Experts language model with 7 billion parameters achieves superior performance by using only 1 billion parameters per input token and outperforms larger models in various experiments.

81sparse Mixture-of-ExpertsOLMoE-1B-7BHF ↗arXiv ↗
12

NVLM: Open Frontier-Class Multimodal LLMs

Wenliang Dai, Nayeon Lee, Boxin Wang +7 authors

NVML 1.0, a family of multimodal large language models, achieves state-of-the-art results in vision-language tasks by combining text-only and multimodal training, utilizing a new architecture and dataset strategy.

75multimodal large language modelsdecoder-only multimodal LLMsHF ↗arXiv ↗
13

Towards a Unified View of Preference Learning for Large Language Models: A Survey

Bofei Gao, Feifan Song, Yibo Miao +21 authors

Large Language Models (LLMs) exhibit remarkably powerful capabilities. One of the crucial factors to achieve success is aligning the LLM's output with human preferences. This alignment process often requires only a small amount of data to efficiently enhance the LLM's performance. While effective, research in this area spans multiple domains, and the methods involved are relatively complex to understand. The relationships between different methods have been under-explored, limiting the development of the preference alignment. In light of this, we break down the existing popular alignment strategies into different components and provide a unified framework to study the current alignment strategies, thereby establishing connections among them. In this survey, we decompose all the strategies in preference learning into four components: model, data, feedback, and algorithm. This unified view offers an in-depth understanding of existing alignment algorithms and also opens up possibilities to synergize the strengths of different strategies. Furthermore, we present detailed working examples of prevalent existing algorithms to facilitate a comprehensive understanding for the readers. Finally, based on our unified perspective, we explore the challenges and future research directions for aligning large language models with human preferences.

72HF ↗arXiv ↗
14

Kvasir-VQA: A Text-Image Pair GI Tract Dataset

Sushant Gautam, Andrea Storås, Cise Midoglu +4 authors

Kvasir-VQA is a dataset with question-and-answer annotations for GI diagnostics, supporting image captioning, VQA, synthetic image generation, object detection, and classification.

71Visual Question Answering (VQA)image captioningHF ↗arXiv ↗
15

Imagine yourself: Tuning-Free Personalized Image Generation

Zecheng He, Bo Sun, Felix Juefei-Xu +14 authors

Imagine yourself is a tuning-free diffusion model for personalized image generation that enhances identity preservation, text alignment, and visual quality through synthetic data generation, parallel attention architecture, and multi-stage fine-tuning.

69diffusion modelspersonalizationHF ↗arXiv ↗
18

Programming Every Example: Lifting Pre-training Data Quality like Experts at Scale

Fan Zhou, Zengzhi Wang, Qian Liu +2 authors

ProX, a novel framework treating data refinement as a programming task, enhances large language model pre-training by generating fine-grained operations for each example, outperforming traditional human-crafted rule-based methods significantly across various benchmarks and domains.

64Programming Every Example (ProX)data refinementHF ↗arXiv ↗
23

MIO: A Foundation Model on Multimodal Tokens

Zekun Wang, King Zhu, Chunpu Xu +14 authors

MIO, a novel foundation model, achieves competitive performance in multimodal tasks through an end-to-end, autoregressive approach using causal multimodal modeling.

53MIOmultimodal tokensHF ↗arXiv ↗
27

MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct

Run Luo, Haonan Zhang, Longze Chen +13 authors

MMEvol, a multimodal instruction data evolution framework, enhances the capabilities of Multimodal Large Language Models by generating diverse and complex image-text instruction datasets, leading to improved performance across various vision-language tasks.

49Multimodal Large Language ModelsMMEvolHF ↗arXiv ↗
29

Prithvi WxC: Foundation Model for Weather and Climate

Johannes Schmude, Sujit Roy, Will Trojak +26 authors

Prithvi WxC, a 2.3 billion parameter foundation model using an encoder-decoder architecture with transformer concepts, addresses weather forecasting, downscaling, and extreme events estimation with a mixed masked reconstruction and forecasting objective.

48encoder-decoder-based architecturetransformer modelsHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号