TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

December 2024
本月最热381

Qwen2.5 Technical Report

Qwen, An Yang, Baosong Yang +39 authors

Qwen2.5, an enhanced series of large language models, demonstrates superior performance across various benchmarks and use cases through extensive pre-training and advanced post-training techniques.

large language modelspre-trainingpost-trainingsupervised finetuningHF ↗arXiv ↗

50 篇论文 · 按点赞排序

03

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Zhe Chen, Weiyun Wang, Yue Cao +37 authors

InternVL 2.5, an advanced multimodal large language model, showcases competitive performance across various benchmarks, including multimodal reasoning and understanding, and is the first open-source model to surpass 70% on the MMMU benchmark using Chain-of-Thought reasoning.

162multimodal large language modelvision encodersHF ↗arXiv ↗
05

PaliGemma 2: A Family of Versatile VLMs for Transfer

Andreas Steiner, André Susano Pinto, Michael Tschannen +15 authors

PaliGemma 2 integrates a SigLIP-So400m vision encoder with Gemma 2 models of varying sizes and resolutions, advancing transfer performance across diverse vision-language tasks, including OCR and captioning.

136SigLIP-So400mvision encoderHF ↗arXiv ↗
06

Phi-4 Technical Report

Marah Abdin, Jyoti Aneja, Harkirat Behl +24 authors

A 14-billion parameter language model surpasses its teacher model in STEM-focused QA through strategic use of synthetic data, improved data quality, and enhanced training techniques.

124training recipedata qualityHF ↗arXiv ↗
09

Byte Latent Transformer: Patches Scale Better Than Tokens

Artidoro Pagnoni, Ram Pasunuru, Pedro Rodriguez +11 authors

A Byte Latent Transformer (BLT) matches tokenization-based LLM performance at scale with improved inference efficiency and robustness, using entropy-based byte patching.

109Byte Latent Transformer (BLT)byte-level LLMHF ↗arXiv ↗
11

GenEx: Generating an Explorable World

Taiming Lu, Tianmin Shu, Junfei Xiao +8 authors

GenEx generates 3D environments from a single image, enabling AI agents to explore and interact with a consistent, expansive space through guided generative imagination.

98Generative imaginationpanoramic video streamsHF ↗arXiv ↗
16

1.58-bit FLUX

Chenglin Yang, Celong Liu, Xueqing Deng +4 authors

1.58-bit quantization of the FLUX.1-dev text-to-image model achieves comparable generation quality with reduced storage and inference overhead using self-supervised methods.

88quantizingself-supervisionHF ↗arXiv ↗
18

ProcessBench: Identifying Process Errors in Mathematical Reasoning

Chujie Zheng, Zhenru Zhang, Beichen Zhang +6 authors

ProcessBench evaluates models' ability to identify errors in mathematical reasoning steps, showing that existing process reward models struggle with difficult problems and underperform compared to critic models and a fine-tuned PRM.

87ProcessBenchprocess reward modelsHF ↗arXiv ↗
21

STIV: Scalable Text and Image Conditioned Video Generation

Zongyu Lin, Wei Liu, Chen Chen +14 authors

STIV, a text-image-conditioned video generation method integrating Diffusion Transformer and classifier-free guidance, achieves state-of-the-art performance in text-to-video, text-image-to-video, and image-to-video tasks.

74video generationmodel architecturesHF ↗arXiv ↗
22

Progressive Multimodal Reasoning via Active Retrieval

Guanting Dong, Chenghao Zhang, Mengjie Deng +3 authors

AR-MCTS enhances multimodal large language models' reasoning capabilities through active retrieval, Monte Carlo Tree Search, and a process reward model, improving performance across multimodal reasoning tasks.

73Active RetrievalMonte Carlo Tree SearchHF ↗arXiv ↗
25

YuLan-Mini: An Open Data-efficient Language Model

Yiwen Hu, Huatong Song, Jia Deng +8 authors

YuLan-Mini, a 2.42B parameter base model, achieves top-tier performance with efficient pre-training techniques, including a data pipeline with cleaning and scheduling, robust optimization, and annealing with targeted data selection.

67data pipelinedata cleaningHF ↗arXiv ↗
29

NVILA: Efficient Frontier Visual Language Models

Zhijian Liu, Ligeng Zhu, Baifeng Shi +24 authors

NVILA, a family of VLMs, optimizes efficiency and accuracy through a scale-then-compress approach, enhancing performance across various benchmarks while reducing computational costs.

62Visual language modelsNVILAHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号