TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Dec 30 – Jan 5, 2025
本周最热110

2.5 Years in Class: A Multimodal Textbook for Vision-Language Pretraining

Wenqi Zhang, Hang Zhang, Xin Li +6 authors

A new high-quality instructional video-based corpus improves VLM pretraining through more coherent context and richer knowledge, enhancing performance in knowledge-intensive tasks.

Vision-Language ModelsVLMsmultimodal textbookinstructional videosHF ↗arXiv ↗

48 篇论文 · 按点赞排序

04

1.58-bit FLUX

Chenglin Yang, Celong Liu, Xueqing Deng +4 authors

1.58-bit quantization of the FLUX.1-dev text-to-image model achieves comparable generation quality with reduced storage and inference overhead using self-supervised methods.

87quantizingself-supervisionHF ↗arXiv ↗
08

LTX-Video: Realtime Video Latent Diffusion

Yoav HaCohen, Nisan Chiprut, Benny Brazowski +13 authors

LTX-Video, a transformer-based latent diffusion model, integrates Video-VAE and denoising transformer for efficient high-resolution video generation with temporal consistency.

51transformer-based latent diffusion modelVideo-VAEHF ↗arXiv ↗
13

Bringing Objects to Life: 4D generation from 3D objects

Ohad Rahamim, Ori Malca, Dvir Samuel +1 authors

A method for animating 3D objects using text prompts to guide 4D generation, employing Neural Radiance Fields and diffusion models, achieves high identity preservation and visual quality.

40Neural Radiance Field (NeRF)Image-to-Video diffusion modelHF ↗arXiv ↗
16

MLLM-as-a-Judge for Image Safety without Human Labeling

Zhenting Wang, Shuming Hu, Shiyu Zhao +12 authors

A method using Multimodal Large Language Models for zero-shot detection of unsafe images improves upon human-labeled fine-tuning by objectifying safety rules and using cascaded reasoning for complex judgments.

31Multimodal Large Language Modelszero-shot settingHF ↗arXiv ↗
17

Xmodel-2 Technical Report

Wang Qun, Liu Yang, Lin Qingquan +2 authors

Xmodel-2, a 1.2-billion-parameter large language model, utilizes a unified set of hyperparameters across different scales and the WSD learning rate scheduler to achieve state-of-the-art performance in complex reasoning tasks with efficient training.

27large language modelreasoning tasksHF ↗arXiv ↗
18

ProgCo: Program Helps Self-Correction of Large Language Models

Xiaoshuai Song, Yanan Wu, Weixun Wang +3 authors

Program-driven Self-Correction (ProgCo) improves large language models' self-verification and refinement by using self-generated and self-executing verification pseudo-programs, particularly in complex reasoning tasks.

26Large language models (LLMs)self-verificationHF ↗arXiv ↗
21

A3: Android Agent Arena for Mobile GUI Agents

Yuxiang Chai, Hanhao Li, Jiayu Zhang +5 authors

Android Agent Arena (A3) provides an evaluation platform for mobile GUI agents with real-world tasks, flexible action spaces, and automated LLM-based testing.

22large language models (LLMs)mobile GUI agentsHF ↗arXiv ↗
25

Edicho: Consistent Image Editing in the Wild

Qingyan Bai, Hao Ouyang, Yinghao Xu +5 authors

Edicho, a training-free diffusion-based editing method, uses explicit image correspondence and a classifier-free guidance denoising strategy for consistent editing across diverse images.

20diffusion modelsattention manipulation moduleHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号