TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

598 篇论文 · 按点赞排序

91

GPT-4o System Card

OpenAI, Aaron Hurst, Adam Lerer +416 authors

GPT-4o is an omnimodal autoregressive model trained to handle text, audio, image, and video inputs, offering high-performance outputs across these modalities, with particular strengths in vision and audio.

88autoregressive modelomnimodalHF ↗arXiv ↗
97

ProcessBench: Identifying Process Errors in Mathematical Reasoning

Chujie Zheng, Zhenru Zhang, Beichen Zhang +6 authors

ProcessBench evaluates models' ability to identify errors in mathematical reasoning steps, showing that existing process reward models struggle with difficult problems and underperform compared to critic models and a fine-tuned PRM.

87ProcessBenchprocess reward modelsHF ↗arXiv ↗
98

Lumiere: A Space-Time Diffusion Model for Video Generation

Omer Bar-Tal, Hila Chefer, Omer Tov +11 authors

A text-to-video diffusion model using Space-Time U-Net architecture generates realistic, diverse, and coherent videos through a single pass, achieving state-of-the-art results and supporting various content creation and editing tasks.

86Space-Time U-Netdiffusion modelHF ↗arXiv ↗
99

OLMo: Accelerating the Science of Language Models

Dirk Groeneveld, Iz Beltagy, Pete Walsh +40 authors

OLMo, an open-sourced language model, provides comprehensive access to training data, code, and architecture, facilitating research and innovation in language modeling.

86open language modellanguage modelsHF ↗arXiv ↗
102

Vision language models are blind

Pooyan Rahmanzadehgervi, Logan Bolton, Mohammad Reza Taesiri +1 authors

State-of-the-art large language models with vision capabilities perform poorly on simple visual tasks, indicating limitations in their visual understanding.

84large language modelsvision capabilitiesHF ↗arXiv ↗
106

Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs

Keen You, Haotian Zhang, Eldon Schoop +5 authors

Ferret-UI, a multimodal large language model tailored for mobile UI screens, enhances understanding and interaction through region annotations and a comprehensive dataset of UI tasks, outperforming existing models including GPT-4V.

83multimodal large language modelsMLLMsHF ↗arXiv ↗
108

The Unreasonable Ineffectiveness of the Deeper Layers

Andrey Gromov, Kushal Tirumala, Hassan Shapourian +2 authors

Layer pruning of pre-trained LLMs with parameter-efficient finetuning methods shows minimal performance degradation and significant resource savings in both finetuning and inference.

82layer-pruningopen-weight pretrained LLMsHF ↗arXiv ↗
111

Better & Faster Large Language Models via Multi-token Prediction

Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière +2 authors

Training large language models to predict multiple future tokens simultaneously enhances their sample efficiency and performance, particularly on generative benchmarks and small algorithmic tasks, with decreased inference time.

82next-token prediction lossmulti-token predictionHF ↗arXiv ↗
114

OLMoE: Open Mixture-of-Experts Language Models

Niklas Muennighoff, Luca Soldaini, Dirk Groeneveld +21 authors

A sparse Mixture-of-Experts language model with 7 billion parameters achieves superior performance by using only 1 billion parameters per input token and outperforms larger models in various experiments.

81sparse Mixture-of-ExpertsOLMoE-1B-7BHF ↗arXiv ↗
4 / 20

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号