TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Sep 25 – Oct 1, 2023
本周最热86

Vision Transformers Need Registers

Timothée Darcet, Maxime Oquab, Julien Mairal +1 authors

Additional input tokens in Vision Transformers mitigate artifacts in feature maps, enhancing performance and enabling smoother visual processing.

Transformersfeature mapsViT networkshigh-norm tokensHF ↗arXiv ↗

43 篇论文 · 按点赞排序

02

CodePlan: Repository-level Coding using LLMs and Planning

Ramakrishna Bairi, Atharv Sonwane, Aditya Kanade +6 authors

CodePlan automates repository-level coding tasks, such as package migration and temporal code edits, using a planning framework that leverages LLMs with context derived from code repositories and change analysis.

80Large Language ModelsLLMsHF ↗arXiv ↗
03

AnyMAL: An Efficient and Scalable Any-Modality Augmented Language Model

Seungwhan Moon, Andrea Madotto, Zhaojiang Lin +10 authors

AnyMAL is a unified model that processes multiple modalities (text, image, video, audio, IMU) and generates text, achieving top performance in multimodal tasks through a pre-trained aligner and fine-tuning with diverse instructions.

56Any-Modality Augmented Language ModelAnyMALHF ↗arXiv ↗
07

Qwen Technical Report

Jinze Bai, Shuai Bai, Yunfei Chu +45 authors

Qwen, a series of large language models, including chat and specialized coding and mathematics variants, exhibit superior performance across various tasks and outperform open-source models.

39LLMslarge language modelsHF ↗arXiv ↗
10

Text-to-3D using Gaussian Splatting

Zilong Chen, Feng Wang, Huaping Liu

Gaussian Splatting based text-to-3D generation improves geometry accuracy and fidelity through progressive optimization, including geometry optimization and appearance refinement.

33Gaussian Splatting3D priorHF ↗arXiv ↗
12

Aligning Large Multimodal Models with Factually Augmented RLHF

Zhiqing Sun, Sheng Shen, Shengcao Cao +9 authors

The paper presents Factually Augmented RLHF to address multimodal misalignment in large multimodal models, significantly improving hallucination reduction and overall performance on vision-language tasks compared to existing methods.

32Reinforcement Learning from Human Feedback (RLHF)vision-language alignmentHF ↗arXiv ↗
13

Effective Long-Context Scaling of Foundation Models

Wenhan Xiong, Jingyu Liu, Igor Molybog +18 authors

Long-context LLMs achieve significant advancements in handling extended contexts through continual pretraining and efficient instruction tuning, surpassing existing models on various benchmarks.

31LLMslong-contextHF ↗arXiv ↗
14

Deep Geometrized Cartoon Line Inbetweening

Li Siyao, Tianpei Gu, Weiye Xiao +3 authors

AnimeInbet addresses line inbetweening by geometrically representing line drawings as graphs and using a vertex correspondence Transformer and visibility predictor for high-quality frame interpolation.

25frame interpolationgraph fusionHF ↗arXiv ↗
15

Finite Scalar Quantization: VQ-VAE Made Simple

Fabian Mentzer, David Minnen, Eirikur Agustsson +1 authors

The proposed finite scalar quantization (FSQ) for VAE latent representations offers competitive performance in image generation and other vision tasks without the complexity of VQ.

24vector quantizationfinite scalar quantizationHF ↗arXiv ↗
20

Demystifying CLIP Data

Hu Xu, Saining Xie, Xiaoqing Ellen Tan +7 authors

MetaCLIP, a metadata-driven data curation approach for language-image pre-training, outperforms CLIP and scales effectively with larger datasets.

20Contrastive Language-Image Pre-trainingCLIPHF ↗arXiv ↗
23

SCREWS: A Modular Framework for Reasoning with Revisions

Kumar Shridhar, Harsh Jhamtani, Hao Fang +3 authors

A modular framework SCREWS for reasoning with revisions in large language models reveals novel strategies and improves performance across various reasoning tasks by enabling selection and heterogeneous approaches.

18SCREWSSamplingHF ↗arXiv ↗
28

Calibrating LLM-Based Evaluator

Yuxuan Liu, Tianchi Yang, Shaohan Huang +6 authors

AutoCalibrate uses an iterative, gradient-free approach to automatically align an off-the-shelf LLM-based evaluator with human preferences, improving its correlation with expert evaluations.

12large language modelsLLMsHF ↗arXiv ↗
30

VidChapters-7M: Video Chapters at Scale

Antoine Yang, Arsha Nagrani, Ivan Laptev +2 authors

The introduction of VidChapters-7M, a large-scale dataset for video chaptering, improves state-of-the-art performance on video-language tasks and demonstrates the benefit of pretraining on extensive datasets.

11video chapter generationvideo chapter groundingHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号