TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

本月最热192

DocLLM: A layout-aware generative language model for multimodal document understanding

Dongsheng Wang, Natraj Raman, Mathieu Sibue +6 authors

DocLLM extends large language models for visual document understanding by focusing on bounding box information and disentangled matrices for multimodal attention, achieving state-of-the-art performance on document intelligence tasks.

multimodal LLMsimage encodersbounding box informationattention mechanismHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

Mixtral of Experts

Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux +23 authors

Mixtral 8x7B, a Sparse Mixture of Experts language model, achieves superior performance across benchmarks by using a selective architecture that leverages fewer active parameters.

162Sparse Mixture of Experts (SMoE)feedforward blocksHF ↗arXiv ↗
03

Self-Rewarding Language Models

Weizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho +3 authors

A study on Self-Rewarding Language Models shows that using LLM-as-a-Judge prompting for iterative DPO training enhances both instruction-following and self-reward generation, leading to superior performance compared to existing systems.

156Self-Rewarding Language ModelsLLM-as-a-JudgeHF ↗arXiv ↗
04

TinyLlama: An Open-Source Small Language Model

Peiyuan Zhang, Guangtao Zeng, Tianduo Wang +1 authors

TinyLlama, a compact 1.1B language model, leverages FlashAttention to achieve high performance in downstream tasks with enhanced computational efficiency.

96FlashAttentionLlama 2HF ↗arXiv ↗
05

Lumiere: A Space-Time Diffusion Model for Video Generation

Omer Bar-Tal, Hila Chefer, Omer Tov +11 authors

A text-to-video diffusion model using Space-Time U-Net architecture generates realistic, diverse, and coherent videos through a single pass, achieving state-of-the-art results and supporting various content creation and editing tasks.

86Space-Time U-Netdiffusion modelHF ↗arXiv ↗
11

TrustLLM: Trustworthiness in Large Language Models

Lichao Sun, Yue Huang, Haoran Wang +64 authors

This study assesses the trustworthiness of large language models across various dimensions, including truthfulness, safety, fairness, robustness, privacy, and machine ethics, finding a positive correlation with utility and highlighting differences between proprietary and open-source models.

69TrustLLMlarge language modelsHF ↗arXiv ↗
14

Understanding LLMs: A Comprehensive Overview from Training to Inference

Yiheng Liu, Hao He, Tianle Han +18 authors

The paper reviews techniques for cost-efficient training and deployment of large language models, covering aspects like data preprocessing, parallel training, model fine-tuning, and inference optimizations including model compression and memory scheduling.

66Large Language Modelspre-training tasksHF ↗arXiv ↗
15

Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data

Lihe Yang, Bingyi Kang, Zilong Huang +3 authors

Depth Anything is a robust monocular depth estimation model built on a large-scale dataset with data augmentation and auxiliary supervision strategies, achieving state-of-the-art results on various datasets and enhancing depth-conditioned ControlNet.

64monocular depth estimationdata engineHF ↗arXiv ↗
18

MambaByte: Token-free Selective State Space Model

Junxiong Wang, Tushaar Gangavarapu, Jing Nathan Yan +1 authors

MambaByte, a token-free byte-level autoregressive model, demonstrates computational efficiency and competitive performance compared to subword token-based models, with the added benefit of fast inference due to linear length scaling.

60MambaBytetoken-free language modelsHF ↗arXiv ↗
24

LLaMA Pro: Progressive LLaMA with Block Expansion

Chengyue Wu, Yukang Gan, Yixiao Ge +5 authors

A new post-pretraining method using expanded Transformer blocks for Large Language Models improves knowledge without catastrophic forgetting, yielding LLaMA Pro-8.3B that excels in general tasks, programming, and mathematics.

54Large Language ModelsLLMsHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号