TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

403 篇论文 · 按点赞排序

152

GLaMM: Pixel Grounding Large Multimodal Model

Hanoona Rasheed, Muhammad Maaz, Sahal Shaji +7 authors

GLaMM is a multimodal model that generates visually grounded language responses with object segmentation masks from both text and optional visual prompts.

36Large Multimodal ModelsLarge Language ModelsHF ↗arXiv ↗
153

Generative Multimodal Models are In-Context Learners

Quan Sun, Yufeng Cui, Xiaosong Zhang +8 authors

A large-scale generative multimodal model with 37 billion parameters demonstrates strong few-shot in-context learning and achieves state-of-the-art performance on multimodal tasks through scaling-up and instruction tuning.

36task-agnostic in-context learninggenerative multimodal modelHF ↗arXiv ↗
154

Learning to Model the World with Language

Jessy Lin, Yuqing Du, Olivia Watkins +4 authors

Dynalang, a multimodal agent, learns to predict future text and image representations using language hints to improve task performance and enrich its understanding.

36multimodal world modelself-supervised learningHF ↗arXiv ↗
158

Fast Segment Anything

Xu Zhao, Wenchao Ding, Yongqi An +5 authors

A speed-up method using a CNN detector with an instance segmentation branch achieves comparable performance to SAM with 50 times faster runtime by training on a small subset of SAM's dataset.

36segment anything modelSAMHF ↗arXiv ↗
161

JudgeLM: Fine-tuned Large Language Models are Scalable Judges

Lianghui Zhu, Xinggang Wang, Xinlong Wang

Large Language Models fine-tuned as scalable judges (JudgeLM) achieve state-of-the-art performance in evaluating open-ended benchmarks through a comprehensive dataset and benchmark, enhancing judgment efficiency and accuracy.

35Large Language Modelsfine-tuningHF ↗arXiv ↗
162

Contrastive Chain-of-Thought Prompting

Yew Ken Chia, Guizhen Chen, Luu Anh Tuan +2 authors

Contrastive chain of thought, utilizing both valid and invalid reasoning examples, improves language model reasoning and generalization compared to conventional methods.

35chain of thoughtreasoningHF ↗arXiv ↗
163

ChatAnything: Facetime Chat with LLM-Enhanced Personas

Yilin Zhao, Xinbin Yuan, Shanghua Gao +4 authors

A framework for generating anthropomorphized personas with diverse voices and appearances from text descriptions using LLMs and generative models, with improved face landmark detection for automatic animation.

35LLM-based charactersin-context learningHF ↗arXiv ↗
165

TeCH: Text-guided Reconstruction of Lifelike Clothed Humans

Yangyi Huang, Hongwei Yi, Yuliang Xiu +4 authors

TeCH reconstructs high-fidelity 3D human models from single images using descriptive text prompts, a personalized Text-to-Image diffusion model, and a hybrid DMTet representation, outperforming existing methods.

35descriptive text promptsgarment parsing modelHF ↗arXiv ↗
168

StarCoder: may the source be with you!

Raymond Li, Loubna Ben Allal, Yangtian Zi +64 authors

StarCoder, a 15.5B parameter LLM trained on 1 trillion tokens, outperforms other open Code LLMs across multiple languages and fine-tuned Python, with safety enhancements and publicly available under the Open Responsible AI Model license.

34Large Language ModelsCode LLMsHF ↗arXiv ↗
171

OtterHD: A High-Resolution Multi-modality Model

Bo Li, Peiyuan Zhang, Jingkang Yang +3 authors

OtterHD-8B, an advanced multimodal model from Fuyu-8B, excels in processing high-resolution inputs and discerning detailed spatial relationships through MagnifierBench, an evaluation framework highlighting the importance of vision encoder flexibility.

34multimodal modelhigh-resolution visual inputsHF ↗arXiv ↗
172

One Wide Feedforward is All You Need

Telmo Pessoa Pires, António V. Lopes, Yannick Assogba +1 authors

The study shows that the Feed Forward Network in Transformers is highly redundant and its reduction or sharing within the model can lead to improved accuracy and latency without significant loss of performance.

34Transformer architectureAttentionHF ↗arXiv ↗
176

LLaSM: Large Language and Speech Model

Yu Shu, Siwei Dong, Guangyao Chen +5 authors

LLaSM, an end-to-end trained large multi-modal speech-language model, enhances human interaction with AI by following speech-and-language instructions using cross-modal conversational abilities.

34multi-modal large language modelsvision-language multi-modal modelsHF ↗arXiv ↗
177

RMT: Retentive Networks Meet Vision Transformers

Qihang Fan, Huaibo Huang, Mingrui Chen +2 authors

The proposed RMT model, combining RetNet and Transformer architectures, introduces explicit spatial distance priors and coordinate-wise decomposition to achieve exceptional performance in computer vision tasks.

34TransformerRetentive Network (RetNet)HF ↗arXiv ↗
180

Seeing the World through Your Eyes

Hadi Alzayer, Kevin Zhang, Brandon Feng +2 authors

A method for reconstructing 3D scenes beyond a camera's line of sight using eye reflections is proposed, refining cornea poses, radiance fields, and iris textures.

34cornea posesradiance fieldHF ↗arXiv ↗
6 / 14

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号