TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

598 篇论文 · 按点赞排序

31

Diffusion Models Are Real-Time Game Engines

Dani Valevski, Yaniv Leviathan, Moab Arar +1 authors

GameNGen, a neural model-powered game engine, simulates high-quality gameplay in real-time using a diffusion model conditioned on past frames and actions.

126neural modelreal-time interactionHF ↗arXiv ↗
33

Phi-4 Technical Report

Marah Abdin, Jyoti Aneja, Harkirat Behl +24 authors

A 14-billion parameter language model surpasses its teacher model in STEM-focused QA through strategic use of synthetic data, improved data quality, and enhanced training techniques.

124training recipedata qualityHF ↗arXiv ↗
35

SAM 2: Segment Anything in Images and Videos

Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu +15 authors

Segment Anything Model 2 (SAM 2) uses a transformer architecture with streaming memory to achieve high performance in image and video segmentation, requiring fewer interactions and faster processing than previous models.

123transformer architecturestreaming memoryHF ↗arXiv ↗
39

The Llama 3 Herd of Models

Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey +530 authors

Llama 3, a multilingual and multi-modal language model with 405B parameters, achieves competitive performance across tasks including image, video, and speech recognition when integrated through a compositional approach.

119TransformermultilingualityHF ↗arXiv ↗
41

Octopus v4: Graph of language models

Wei Chen, Zhiyuan Li

The Octopus v4 model uses functional tokens to integrate and direct queries to task-specific open-source language models, achieving SOTA performance with models under 10B parameters.

117functional tokensOctopus v4HF ↗arXiv ↗
43

KAN: Kolmogorov-Arnold Networks

Ziming Liu, Yixuan Wang, Sachin Vaidya +5 authors

Kolmogorov-Arnold Networks (KANs) outperform Multi-Layer Perceptrons (MLPs) in accuracy and interpretability by using learnable activation functions and spline-based weights.

116Kolmogorov-Arnold NetworksKANsHF ↗arXiv ↗
44

OmniGen: Unified Image Generation

Shitao Xiao, Yueze Wang, Junjie Zhou +6 authors

OmniGen is a unified diffusion model for image generation that supports diverse tasks without additional modules, emphasizing simplicity, knowledge transfer, and reasoning capabilities.

115diffusion modelOmniGenHF ↗arXiv ↗
47

Chain-of-Thought Reasoning Without Prompting

Xuezhi Wang, Denny Zhou

LLMs can perform chain-of-thought reasoning through top-k decoding without manual prompt engineering, outperforming greedy decoding and showing higher confidence in answers.

111large language models (LLMs)chain-of-thought (CoT) promptingHF ↗arXiv ↗
48

Aria: An Open Multimodal Native Mixture-of-Experts Model

Dongxu Li, Yudong Liu, Haoning Wu +7 authors

Aria is an open multimodal native AI model with best-in-class performance across various tasks, designed with a mixture-of-experts architecture and pre-trained through a four-stage pipeline.

111mixture-of-expert modelvisual tokenHF ↗arXiv ↗
50

Byte Latent Transformer: Patches Scale Better Than Tokens

Artidoro Pagnoni, Ram Pasunuru, Pedro Rodriguez +11 authors

A Byte Latent Transformer (BLT) matches tokenization-based LLM performance at scale with improved inference efficiency and robustness, using entropy-based byte patching.

109Byte Latent Transformer (BLT)byte-level LLMHF ↗arXiv ↗
55

Depth Anything V2

Lihe Yang, Bingyi Kang, Zilong Huang +4 authors

Depth Anything V2 improves monocular depth estimation through synthetic images, larger teacher models, and pseudo-labeled real images, achieving better efficiency and accuracy than Stable Diffusion models.

105monocular depth estimationsynthetic imagesHF ↗arXiv ↗
56

What matters when building vision-language models?

Hugo Laurençon, Léo Tronchon, Matthieu Cord +1 authors

Idefics2, a vision-language model with 8 billion parameters, achieves state-of-the-art performance on multimodal benchmarks through extensive experimental validation.

104vision-language modelslarge language modelsHF ↗arXiv ↗
58

ReFT: Representation Finetuning for Language Models

Zhengxuan Wu, Aryaman Arora, Zheng Wang +4 authors

Representation Finetuning (ReFT) methods, exemplified by Low-rank Linear Subspace ReFT (LoReFT), achieve high efficiency and performance by adapting representations in frozen base models, outperforming state-of-the-art Parameter-efficient Fine-tuning (PEFT) methods.

101Parameter-efficient fine-tuningRepresentation FinetuningHF ↗arXiv ↗
60

Neural Network Diffusion

Kai Wang, Zhaopan Xu, Yukun Zhou +4 authors

Diffusion models can generate high-performing neural network parameters using an autoencoder, producing new subsets of network parameters with comparable or improved performance.

100diffusion modelsautoencoderHF ↗arXiv ↗
2 / 20

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号