TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

December 2023

50 篇论文 · 按点赞排序

32

CogAgent: A Visual Language Model for GUI Agents

Wenyi Hong, Weihan Wang, Qingsong Lv +8 authors

CogAgent, a visual language model with strong GUI understanding and navigation capabilities, outperforms LLM-based methods in both PC and Android GUI tasks using only screenshots.

32visual language modelGUI understandingHF ↗arXiv ↗
33

Relightable Gaussian Codec Avatars

Shunsuke Saito, Gabriel Schwartz, Tomas Simon +2 authors

Relightable Gaussian Codec Avatars model high-fidelity head avatars with real-time relighting capabilities using 3D Gaussians for geometry and learnable radiance transfer for appearance.

323D Gaussiansrelightable appearance modelHF ↗arXiv ↗
34

FaceStudio: Put Your Face Everywhere in Seconds

Yuxuan Yan, Chi Zhang, Rui Wang +3 authors

A hybrid guidance framework for identity-preserving image synthesis efficiently generates stylistic portraits by combining stylized images, facial images, and textual prompts.

32identity-preserving synthesisTextual InversionHF ↗arXiv ↗
41

DreamTuner: Single Image is Enough for Subject-Driven Generation

Miao Hua, Jiawei Liu, Fei Ding +3 authors

DreamTurner uses a novel approach by injecting reference information through subject encoders and self-subject-attention layers to enhance subject-driven image generation, balancing subject learning and model capabilities.

27diffusion-based modelstext-to-image generationHF ↗arXiv ↗
42

Zero-Shot Metric Depth with a Field-of-View Conditioned Diffusion Model

Saurabh Saxena, Junhwa Hur, Charles Herrmann +2 authors

A generic diffusion model with log-scale depth parameterization and FOV conditioning achieves state-of-the-art zero-shot metric depth estimation by handling indoor and outdoor scenes effectively and reducing relative error significantly.

27monocular depth estimationdiffusion modelHF ↗arXiv ↗
45

FreeInit: Bridging Initialization Gap in Video Diffusion Models

Tianxing Wu, Chenyang Si, Yuming Jiang +2 authors

FreeInit addresses the temporal consistency and unnatural dynamics issues in diffusion-based video generation by refining spatial-temporal low-frequency components during inference, improving subject appearance and consistency.

27diffusion-based video generationtemporal consistencyHF ↗arXiv ↗
46

Gaussian Head Avatar: Ultra High-fidelity Head Avatar via Dynamic Gaussians

Yuelang Xu, Benwang Chen, Zhe Li +4 authors

The method uses controllable 3D Gaussians and a fully learned MLP-based deformation field for high-fidelity 3D head avatar modeling under sparse views, employing geometry-guided initialization with implicit SDF and Deep Marching Tetrahedra for stability and high rendering quality.

273D GaussiansMLP-based deformation fieldHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号