TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

50 篇论文 · 按点赞排序

31

DeepSeek-VL: Towards Real-World Vision-Language Understanding

Haoyu Lu, Wen Liu, Bo Zhang +11 authors

DeepSeek-VL is an open-source vision-language model that achieves state-of-the-art performance in real-world applications by combining a hybrid vision encoder with effective pretraining strategies to preserve language model capabilities.

50DeepSeek-VLVision-Language ModelHF ↗arXiv ↗
35

Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM

Sainbayar Sukhbaatar, Olga Golovneva, Vasu Sharma +8 authors

Branch-Train-MiX (BTX) method enhances Large Language Models by asynchronously training experts in parallel and integrating them using Mixture-of-Expert layers with token-level routing for improved accuracy and efficiency.

45Large Language ModelsBranch-Train-MiXHF ↗arXiv ↗
36

ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Xiwei Hu, Rui Wang, Yixiao Fang +3 authors

ELLA, an Efficient Large Language Model Adapter, enhances text-to-image diffusion models by integrating powerful Large Language Models through a Timestep-Aware Semantic Connector, improving dense prompt comprehension and generation quality.

45diffusion modelstext-to-image generationHF ↗arXiv ↗
37

sDPO: Don't Use Your Data All at Once

Dahyun Kim, Yungi Kim, Wonho Song +4 authors

A stepwise direct preference optimization approach improves the alignment of large language models with human preferences and enhances their performance.

41large language modelsLLMHF ↗arXiv ↗
43

VLOGGER: Multimodal Diffusion for Embodied Avatar Synthesis

Enric Corona, Andrei Zanfir, Eduard Gabriel Bazavan +3 authors

VLOGGER generates audio-driven human videos from a single image using a diffusion-based method that includes 3D motion and text-to-image models, outperforming existing methods in quality, identity, and consistency.

36stochastic human-to-3d-motion diffusion modeldiffusion-based architectureHF ↗arXiv ↗
44

InternLM2 Technical Report

Zheng Cai, Maosong Cao, Haojiong Chen +97 authors

InternLM2 is an open-source LLM that outperforms predecessors through innovative pre-training and optimization techniques, including Supervised Fine-Tuning and Conditional Online Reinforcement Learning from Human Feedback.

35Large Language ModelsLLMsHF ↗arXiv ↗
49

Can large language models explore in-context?

Akshay Krishnamurthy, Keegan Harris, Dylan J. Foster +2 authors

Experimentation with LLMs in multi-armed bandit environments reveals that robust exploration requires specific prompts, external summarization of history, or algorithmic interventions.

32Large Language ModelsLLMsHF ↗arXiv ↗
50

mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Anwen Hu, Haiyang Xu, Jiabo Ye +8 authors

A Unified Structure Learning framework is introduced to enhance Visual Document Understanding by integrating structure-aware parsing and multi-grained text localization tasks using a vision-to-text module H-Reducer, achieving state-of-the-art results on multiple benchmarks.

32Multimodal Large Language ModelsVisual Document UnderstandingHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号