TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jan 13 – Jan 19, 2025
本周最热304

MiniMax-01: Scaling Foundation Models with Lightning Attention

MiniMax, Aonian Li, Bangwei Gong +87 authors

The MiniMax-01 series, including MiniMax-Text-01 and MiniMax-VL-01, offer superior long-context processing and match state-of-the-art model performance with longer context windows through lightning attention, Mixture of Experts (MoE), and efficient parallel strategies.

lightning attentionMixture of ExpertsMoEparallel strategyHF ↗arXiv ↗

50 篇论文 · 按点赞排序

03

Tensor Product Attention Is All You Need

Yifan Zhang, Yifeng Liu, Huizhuo Yuan +4 authors

T6, a new Transformer architecture using Tensor Product Attention, improves language model performance and memory efficiency, allowing for longer sequence processing.

90Tensor Product AttentionTPAHF ↗arXiv ↗
06

LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs

Omkar Thawakar, Dinura Dissanayake, Ketan More +12 authors

A framework for evaluating and improving step-by-step visual reasoning in large language models using a specialized benchmark and a novel multimodal model trained with curriculum learning.

67visual reasoninglarge language modelsHF ↗arXiv ↗
07

Enabling Scalable Oversight via Self-Evolving Critic

Zhengyang Tang, Ziniu Li, Zhenyang Xiao +8 authors

SCRIT, a self-evolving critique framework, enhances LLMs' critique capabilities using synthetic data and self-validation, achieving significant improvements in critique-correction and error identification benchmarks.

65Large Language Models (LLMs)self-evolvingHF ↗arXiv ↗
09

Towards Best Practices for Open Datasets for LLM Training

Stefan Baack, Stella Biderman, Kasia Odrozek +36 authors

Many AI companies are training their large language models (LLMs) on data without the permission of the copyright owners. The permissibility of doing so varies by jurisdiction: in countries like the EU and Japan, this is allowed under certain restrictions, while in the United States, the legal landscape is more ambiguous. Regardless of the legal status, concerns from creative producers have led to several high-profile copyright lawsuits, and the threat of litigation is commonly cited as a reason for the recent trend towards minimizing the information shared about training datasets by both corporate and public interest actors. This trend in limiting data information causes harm by hindering transparency, accountability, and innovation in the broader ecosystem by denying researchers, auditors, and impacted individuals access to the information needed to understand AI models. While this could be mitigated by training language models on open access and public domain data, at the time of writing, there are no such models (trained at a meaningful scale) due to the substantial technical and sociological challenges in assembling the necessary corpus. These challenges include incomplete and unreliable metadata, the cost and complexity of digitizing physical records, and the diverse set of legal and technical skills required to ensure relevance and responsibility in a quickly changing landscape. Building towards a future where AI systems can be trained on openly licensed data that is responsibly curated and governed requires collaboration across legal, technical, and policy domains, along with investments in metadata standards, digitization, and fostering a culture of openness.

61HF ↗arXiv ↗
11

Transformer^2: Self-adaptive LLMs

Qi Sun, Edoardo Cetin, Yujin Tang

A self-adaptive framework for large language models uses reinforcement learning to dynamically adjust task-specific components during inference, enhancing adaptability and performance with efficiency.

55self-adaptive large language models (LLMs)fine-tuningHF ↗arXiv ↗
13

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Qian Chen, Yafeng Chen, Yanni Chen +33 authors

MinMo, a multimodal large language model, integrates speech and text processing to achieve state-of-the-art performance in voice comprehension and generation, while enabling full-duplex conversation and instruction-following capabilities.

54multimodal large language modelnative modelsHF ↗arXiv ↗
24

VideoAuteur: Towards Long Narrative Video Generation

Junfei Xiao, Feng Cheng, Lu Qi +5 authors

A large-scale cooking video dataset and a Long Narrative Video Director model improve visual and semantic coherence in video generation by aligning text and image embeddings.

31Vision-Language Modelsvideo generation modelsHF ↗arXiv ↗
25

MMDocIR: Benchmarking Multi-Modal Retrieval for Long Documents

Kuicai Dong, Yujing Chang, Xin Deik Goh +3 authors

A new benchmark, MMDocIR, is introduced for evaluating multi-modal document retrieval systems, highlighting the superior performance of visual retrievers over text-only methods and emphasizing the importance of visual elements.

29multi-modal document retrievalpage-level retrievalHF ↗arXiv ↗
27

FAST: Efficient Action Tokenization for Vision-Language-Action Models

Karl Pertsch, Kyle Stachowicz, Brian Ichter +6 authors

A new tokenization method called FAST, based on the discrete cosine transform, enables training of autoregressive vision-language-action policies for high-frequency robotic tasks with improved performance and reduced training time.

28autoregressive sequence modelsTransformer-based vision-language action policiesHF ↗arXiv ↗
29

WebWalker: Benchmarking LLMs in Web Traversal

Jialong Wu, Wenbiao Yin, Yong Jiang +6 authors

WebWalkerQA assesses LLMs' ability to traverse websites for high-quality data, showing enhancements when combined with RAG using the WebWalker multi-agent framework.

23Retrieval-augmented generationRAGHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号