TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

本月最热175

Transformer Explainer: Interactive Learning of Text-Generative Models

Aeree Cho, Grace C. Kim, Alexander Karpekov +5 authors

Transformer Explainer is an interactive visualization tool that allows non-experts to understand the inner workings of the GPT-2 model through real-time experimentation and visualization in a web browser.

TransformersGPT-2interactive visualizationmodel overviewHF ↗arXiv ↗

50 篇论文 · 按点赞排序

04

Diffusion Models Are Real-Time Game Engines

Dani Valevski, Yaniv Leviathan, Moab Arar +1 authors

GameNGen, a neural model-powered game engine, simulates high-quality gameplay in real-time using a diffusion model conditioned on past frames and actions.

127neural modelreal-time interactionHF ↗arXiv ↗
06

SAM 2: Segment Anything in Images and Videos

Nikhila Ravi, Valentin Gabeur, Yuan-Ting Hu +15 authors

Segment Anything Model 2 (SAM 2) uses a transformer architecture with streaming memory to achieve high performance in image and video segmentation, requiring fewer interactions and faster processing than previous models.

123transformer architecturestreaming memoryHF ↗arXiv ↗
07

The Llama 3 Herd of Models

Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey +530 authors

Llama 3, a multilingual and multi-modal language model with 405B parameters, achieves competitive performance across tasks including image, video, and speech recognition when integrated through a compositional approach.

119TransformermultilingualityHF ↗arXiv ↗
09

MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Yuan Yao, Tianyu Yu, Ao Zhang +20 authors

MiniCPM-V presents a series of efficient Multimodal Large Language Models optimized for end-side deployment, offering high performance and practical usability compared to larger models.

96Multimodal Large Language ModelsMLLMsHF ↗arXiv ↗
10

Law of Vision Representation in MLLMs

Shijia Yang, Bohan Zhai, Quanzeng You +3 authors

Correlation between cross-modal alignment and vision representation improves performance in multimodal large language models, enabling identification and training of optimal vision representation with reduced computational cost.

95cross-modal alignmentvision representationHF ↗arXiv ↗
11

Sapiens: Foundation for Human Vision Models

Rawal Khirodkar, Timur Bagautdinov, Julieta Martinez +5 authors

Sapiens, a family of vision models for human-centric tasks, achieves superior performance with self-supervised pretraining and can be easily fine-tuned for 2D pose estimation, body-part segmentation, depth estimation, and surface normal prediction.

93self-supervised pretrainingmodel adaptationHF ↗arXiv ↗
14

Gemma 2: Improving Open Language Models at a Practical Size

Gemma Team, Morgane Riviere, Shreya Pathak +193 authors

Gemma 2 introduces improvements in the Transformer architecture through interleaving local-global attentions and group-query attention, showcasing superior performance relative to its size.

80Transformer architectureinterleaving local-global attentionsHF ↗arXiv ↗
18

Controllable Text Generation for Large Language Models: A Survey

Xun Liang, Hanyu Wang, Yezhaohui Wang +8 authors

Controllable Text Generation techniques for Large Language Models ensure predefined control conditions and high-quality text output, covering content and attribute control through various methods like retraining, fine-tuning, and latent manipulation.

65Large Language ModelsControllable Text GenerationHF ↗arXiv ↗
24

Imagen 3

Imagen-Team-Google, Jason Baldridge, Jakob Bauer +248 authors

Imagen 3, a latent diffusion model, generates high-quality images from text prompts and outperforms state-of-the-art models while addressing safety and representation issues.

62latent diffusion modelHF ↗arXiv ↗
25

LLaVA-OneVision: Easy Visual Task Transfer

Bo Li, Yuanhan Zhang, Dong Guo +7 authors

LLaVA-OneVision is a unified multimodal model that advances performance across single-image, multi-image, and video scenarios with strong transfer learning capabilities.

61multimodal modelsLMMsHF ↗arXiv ↗
26

Hermes 3 Technical Report

Ryan Teknium, Jeffrey Quesnelle, Chen Guang

Hermes 3, a neutrally-aligned instruct and tool use model with strong reasoning and creative capabilities, achieves top performance on public benchmarks.

60instruct-tuned modelslarge language modelsHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号