TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Nov 25 – Dec 1, 2024

50 篇论文 · 按点赞排序

32

Factorized Visual Tokenization and Generation

Zechen Bai, Jianxiong Gao, Ziteng Gao +4 authors

Factorized Quantization (FQ) improves image generation by decomposing large codebooks into smaller, diverse sub-codebooks, enhancing scalability and reconstruction quality through disentanglement regularization and semantic representation learning.

18Visual tokenizerstransformer-based modelsHF ↗arXiv ↗
35

BALROG: Benchmarking Agentic LLM and VLM Reasoning On Games

Davide Paglieri, Bartłomiej Cupiał, Samuel Coward +10 authors

BALROG, a novel benchmark, evaluates the agentic capabilities of large language models and vision language models across diverse, challenging games with fine-grained metrics, revealing significant performance gaps, especially in vision-based decision-making.

18Large Language ModelsVision Language ModelsHF ↗arXiv ↗
36

Diffusion Self-Distillation for Zero-Shot Customized Image Generation

Shengqu Cai, Eric Chan, Yunzhi Zhang +3 authors

Diffusion Self-Distillation uses a pre-trained text-to-image model to generate a dataset for fine-tuning into a text+image-to-image model, improving identity-preservation generation without test-time optimization.

16text-to-image diffusion modelsidentity-preserving generationHF ↗arXiv ↗
37

TEXGen: a Generative Diffusion Model for Mesh Textures

Xin Yu, Ze Yuan, Yuan-Chen Guo +6 authors

A large diffusion model trained directly in UV texture space generates high-resolution texture maps from text prompts and images, supporting applications like inpainting, completion, and synthesis.

16diffusion modelsUV texture spaceHF ↗arXiv ↗
43

VisualLens: Personalization through Visual History

Wang Bill Zhu, Deqing Fu, Kai Sun +8 authors

VisualLens extracts, filters, and refines image representations from users' visual histories to enhance personalization, outperforming existing methods and GPT-4o in recommendation tasks.

15image representationsvisual historiesHF ↗arXiv ↗
48

LongKey: Keyphrase Extraction for Long Documents

Jeovane Honorio Alves, Radu State, Cinthia Obladen de Almendra Freitas +1 authors

LongKey, a novel framework using an encoder-based language model and max-pooling embedder, extracts keyphrases from lengthy documents, outperforming existing methods on comprehensive datasets.

12encoder-based language modelmax-pooling embedderHF ↗arXiv ↗
50

Collaborative Decoding Makes Visual Auto-Regressive Modeling Efficient

Zigeng Chen, Xinyin Ma, Gongfan Fang +1 authors

Collaborative Decoding (CoDe) enhances Visual Auto-Regressive (VAR) modeling efficiency by integrating a large and small model to reduce memory usage and computation time without significantly affecting image quality.

12Visual Auto-Regressive (VAR) modelingCollaborative Decoding (CoDe)HF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号