TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Nov 11 – Nov 17, 2024
本周最热80

MagicQuill: An Intelligent Interactive Image Editing System

Zichen Liu, Yue Yu, Hao Ouyang +6 authors

MagicQuill integrates an interface with a multimodal large language model and diffusion prior to enable efficient and precise real-time image editing through minimal user input.

multimodal large language modeldiffusion priortwo-branch plug-in moduleHF ↗arXiv ↗

48 篇论文 · 按点赞排序

02

LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models

Zhengyi Wang, Jonathan Lorraine, Yikai Wang +4 authors

The work demonstrates the capability of LLMs to generate 3D meshes from text by introducing a novel approach to tokenize 3D mesh data, allowing the unification of 3D and text modalities without expanding the model's vocabulary.

78large language modelsLLMsHF ↗arXiv ↗
05

Cut Your Losses in Large-Vocabulary Language Models

Erik Wijmans, Brody Huval, Alexander Hertzberg +2 authors

Cut Cross-Entropy reduces the memory footprint of large language models during training by efficiently computing cross-entropy loss without materializing logits for all tokens.

50Cross-EntropyCut Cross-Entropy (CCE)HF ↗arXiv ↗
08

Stronger Models are NOT Stronger Teachers for Instruction Tuning

Zhangchen Xu, Fengqing Jiang, Luyao Niu +2 authors

The Larger Models' Paradox reveals that larger models are not always better teachers for fine-tuning smaller models, and a new metric, Compatibility-Adjusted Reward (CAR), is introduced to measure and improve the effectiveness of response generators.

39instruction tuninglarge language models (LLMs)HF ↗arXiv ↗
14

SAMPart3D: Segment Any Part in 3D Objects

Yunhan Yang, Yukun Huang, Yuan-Chen Guo +5 authors

SAMPart3D is a scalable and flexible zero-shot 3D part segmentation framework that leverages text-agnostic vision foundation models and multi-view renderings for semantic labeling.

29Vision Language ModelsVLMsHF ↗arXiv ↗
17

Watermark Anything with Localized Messages

Tom Sander, Pierre Fernandez, Alain Durmus +2 authors

The Watermark Anything Model (WAM) is a deep-learning approach for localized image watermarking, offering high imperceptibility and robustness, including capabilities to locate and extract distinct watermarks from small image regions.

22deep-learninglocalized image watermarkingHF ↗arXiv ↗
20

Autoregressive Models in Vision: A Survey

Jing Xiong, Gongye Liu, Lun Huang +17 authors

This survey explores the application of autoregressive models in computer vision, categorizing them into pixel-based, token-based, and scale-based models, and discussing their integration with other generative models and applications in emerging domains.

19autoregressive modelingpixel-levelHF ↗arXiv ↗
21

Balancing Pipeline Parallelism with Vocabulary Parallelism

Man Tsung Yeung, Penghui Qi, Min Lin +1 authors

Proposed techniques achieve balanced computation and memory usage in pipeline parallelism for transformer-based models by partitioning vocabulary layers and optimizing communication barriers, resulting in improved throughput and reduced peak memory usage.

19pipeline parallelismtransformer-based modelsHF ↗arXiv ↗
23

Direct Preference Optimization Using Sparse Feature-Level Constraints

Qingyu Yin, Chak Tou Leong, Hongbo Zhang +8 authors

A novel method, Feature-level constrained Preference Optimization (FPO), improves the alignment of large language models with human preferences by using sparse features from Sparse Autoencoders and feature-level offline reference, achieving better efficiency and win rate.

17Large language modelsReinforcement Learning from Human Feedback (RLHF)HF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号