TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Feb 17 – Feb 23, 2025
本周最热218

Qwen2.5-VL Technical Report

Shuai Bai, Keqin Chen, Xuejing Liu +24 authors

Qwen2.5-VL, the latest vision-language model, advances visual recognition, document parsing, and video comprehension through dynamic resolution processing, Window Attention, and a native Vision Transformer.

Vision TransformerWindow Attentiondynamic resolution processingbounding boxesHF ↗arXiv ↗

50 篇论文 · 按点赞排序

03

SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Michael Tschannen, Alexey Gritsenko, Xiao Wang +11 authors

SigLIP 2, a multilingual vision-language encoder, improves upon SigLIP with unified training techniques, enhancing performance in zero-shot classification, image-text retrieval, localization, and dense prediction across various model sizes and data diversity.

169vision-language encoderscaptioning-based pretrainingHF ↗arXiv ↗
05

Large Language Diffusion Models

Shen Nie, Fengqi Zhu, Zebin You +7 authors

LLaDA, a diffusion model trained from scratch, outperforms autoregressive models in benchmarks and demonstrates strong instruction-following capabilities, challenging the dominance of ARMs in LLMs.

129autoregressive modelsLLaDAHF ↗arXiv ↗
08

Soundwave: Less is More for Speech-Text Alignment in LLMs

Yuhao Zhang, Zhiheng Liu, Fan Bu +3 authors

Soundwave addresses the representation space gap and sequence length inconsistency in end-to-end speech large language models using an efficient training strategy and novel architecture, outperforming Qwen2-Audio with significantly less data.

78large language modelslarge-scale annotated dataHF ↗arXiv ↗
10

S*: Test Time Scaling for Code Generation

Dacheng Li, Shiyi Cao, Chengkun Cao +6 authors

A hybrid test-time scaling framework improves code generation coverage and accuracy across various models and domains.

63hybrid test-time scaling frameworkparallel scalingHF ↗arXiv ↗
11

Magma: A Foundation Model for Multimodal AI Agents

Jianwei Yang, Reuben Tan, Qianhui Wu +10 authors

Magma is a multimodal foundation model with both verbal intelligence and spatial-temporal intelligence, trained on diverse datasets to perform agentic tasks like UI navigation and robotic manipulation, outperforming specialized models.

58vision-language modelsspatial-temporal intelligenceHF ↗arXiv ↗
14

Continuous Diffusion Model for Language Modeling

Jaehyeong Jo, Sung Ju Hwang

A continuous diffusion model for language modeling that leverages the geometry of discrete distributions outperforms existing discrete models and matches autoregressive models in performance.

53diffusion modelsautoregressive modelsHF ↗arXiv ↗
15

Region-Adaptive Sampling for Diffusion Transformers

Ziming Liu, Yifan Yang, Chengruidong Zhang +4 authors

RAS, a novel sampling strategy for diffusion transformers, dynamically adjusts sampling ratios based on regions of focus, achieving speedups in diffusion models with minimal quality loss.

53diffusion modelssampling strategyHF ↗arXiv ↗
16

MMTEB: Massive Multilingual Text Embedding Benchmark

Kenneth Enevoldsen, Isaac Chung, Imene Kerboua +83 authors

The MMTEB benchmark expands MTEB with over 500 tasks across 250+ languages to comprehensively evaluate text embeddings, finding that smaller multilingual models can outperform large LLMs in many cases.

48text embeddingsMassive Multilingual Text Embedding Benchmark (MMTEB)HF ↗arXiv ↗
18

SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?

Samuel Miserendino, Michele Wang, Tejal Patwardhan +1 authors

We introduce SWE-Lancer, a benchmark of over 1,400 freelance software engineering tasks from Upwork, valued at \1 million USD total in real-world payouts. SWE-Lancer encompasses both independent engineering tasks--ranging from 50 bug fixes to \$32,000 feature implementations--and managerial tasks, where models choose between technical implementation proposals. Independent tasks are graded with end-to-end tests triple-verified by experienced software engineers, while managerial decisions are assessed against the choices of the original hired engineering managers. We evaluate model performance and find that frontier models are still unable to solve the majority of tasks. To facilitate future research, we open-source a unified Docker image and a public evaluation split, SWE-Lancer Diamond (https://github.com/openai/SWELancer-Benchmark). By mapping model performance to monetary value, we hope SWE-Lancer enables greater research into the economic impact of AI model development.

46HF ↗arXiv ↗
30

MoM: Linear Sequence Modeling with Mixture-of-Memories

Jusen Du, Weigao Sun, Disen Lan +2 authors

Mixture-of-Memories (MoM) uses independent memory states to enhance recall in sequence modeling while maintaining computational efficiency, outperforming linear models and achieving performance comparable to Transformers.

36linear attentionstate space modelingHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号