TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jan 29 – Feb 4, 2024

50 篇论文 · 按点赞排序

32

Anything in Any Scene: Photorealistic Video Object Insertion

Chen Bai, Zeman Shao, Guoxiang Zhang +11 authors

A novel framework for realistic video simulation, AnyObject in Any Scene, achieves high geometric, lighting, and photorealism by integrating objects, simulating lighting and shadows, and refining video output through style transfer.

17video simulationvirtual realityHF ↗arXiv ↗
34

Transfer Learning for Text Diffusion Models

Kehang Han, Kathleen Kenealy, Aditya Barua +2 authors

Text diffusion models show promise in code synthesis and extractive QA, outperforming autoregressive models in some cases and offering faster long text generation through lightweight adaptation.

17text diffusionautoregressive (AR) decodingHF ↗arXiv ↗
36

Weak-to-Strong Jailbreaking on Large Language Models

Xuandong Zhao, Xianjun Yang, Tianyu Pang +4 authors

A weak-to-strong jailbreaking attack uses smaller, less aligned language models to exploit larger, aligned models, highlighting a significant security concern in language model alignment.

16language modelsLLMsHF ↗arXiv ↗
37

Machine Unlearning for Image-to-Image Generative Models

Guihong Li, Hsiang Hsu, Chun-Fu +2 authors

A novel framework and algorithm for machine unlearning in image-to-image generative models is introduced, showing minimal performance loss and compliance with data retention policies.

15machine unlearningimage-to-image generative modelsHF ↗arXiv ↗
38

High-Quality Image Restoration Following Human Instructions

Marcos V. Conde, Gregor Geigle, Radu Timofte

InstructIR is an all-in-one image restoration method that leverages human-written instructions to guide the restoration process, achieving state-of-the-art results across multiple tasks.

14All-In-One image restorationnatural language promptsHF ↗arXiv ↗
40

Repositioning the Subject within Image

Yikai Wang, Chenjie Cao, Qiaole Dong +2 authors

A diffusion generative model, using SEELE framework, handles dynamic subject repositioning by reformulating the task into unified prompt-guided inpainting with task inversion techniques.

14diffusion generative modelsubject repositioningHF ↗arXiv ↗
44

AToM: Amortized Text-to-Mesh using 2D Diffusion

Guocheng Qian, Junli Cao, Aliaksandr Siarohin +12 authors

Amortized Text-to-Mesh (AToM) generates high-quality textured meshes efficiently from text prompts using a novel triplane-based architecture and achieves superior accuracy and generalization compared to existing methods.

11feed-forward text-to-meshtriplane-based architectureHF ↗arXiv ↗
48

CARFF: Conditional Auto-encoded Radiance Field for 3D Scene Forecasting

Jiezhi Yang, Khushi Desai, Charles Packer +4 authors

A method for predicting future 3D scenes using a conditional auto-encoded radiance field, mapping images to plausible 3D latent configurations and leveraging probabilistic and auto-regressive predictions for autonomous driving applications.

9Conditional Auto-encoded Radiance FieldCARFFHF ↗arXiv ↗
49

MouSi: Poly-Visual-Expert Vision-Language Models

Xiaoran Fan, Tao Ji, Changhao Jiang +21 authors

The ensemble experts technique with a fusion network improves VLM performance by combining visual encoders and optimizing positional encoding for lengthy image feature sequences.

9ensemble expertsvisual encodersHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号