TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Aug 28 – Sep 3, 2023
本周最热106

MVDream: Multi-view Diffusion for 3D Generation

Yichun Shi, Peng Wang, Jianglong Ye +3 authors

MVDream generates geometrically consistent multi-view images from text prompts using pre-trained image diffusion models and Score Distillation Sampling, improving 3D generation stability and supporting personalized generation.

multi-view diffusion modeltext promptimage diffusion modelslarge-scale web datasetsHF ↗arXiv ↗

22 篇论文 · 按点赞排序

04

LLaSM: Large Language and Speech Model

Yu Shu, Siwei Dong, Guangyao Chen +5 authors

LLaSM, an end-to-end trained large multi-modal speech-language model, enhances human interaction with AI by following speech-and-language instructions using cross-modal conversational abilities.

34multi-modal large language modelsvision-language multi-modal modelsHF ↗arXiv ↗
08

Emergence of Segmentation with Minimalistic White-Box Transformers

Yaodong Yu, Tianzhe Chu, Shengbang Tong +4 authors

Transformers designed with an emphasis on low-dimensional structures, such as CRATE, can develop segmentation properties even with minimal supervised training, challenging the notion that such properties require intricate self-supervised mechanisms.

17vision transformersViTsHF ↗arXiv ↗
09

SoTaNa: The Open-Source Software Development Assistant

Ensheng Shi, Fengji Zhang, Yanlin Wang +6 authors

SoTaNa utilizes ChatGPT-generated data and parameter-efficient fine-tuning to enhance LLaMA, providing an accessible open-source software development assistant capable of answering Stack Overflow questions, code summarization, and generation.

14large language modelsChatGPTHF ↗arXiv ↗
14

Active Neural Mapping

Zike Yan, Haoxiang Yang, Hongbin Zha

The paper introduces Active Neural Mapping, a method using a coordinate-based implicit neural representation to guide an agent for active exploration and mapping in unseen environments, leveraging the uncertainty in the neural field weights and geometric information.

11active mappingneural scene representationHF ↗arXiv ↗
20

Relighting Neural Radiance Fields with Shadow and Highlight Hints

Chong Zeng, Guojun Chen, Yue Dong +3 authors

A new neural implicit radiance representation models object relighting from a small set of photographs using signed distance functions and two MLPs, incorporating shadow and highlight hints for high-frequency light transport effects.

9neural implicit radiance representationsigned distance functionHF ↗arXiv ↗
21

Learning Vision-based Pursuit-Evasion Robot Policies

Andrea Bajcsy, Antonio Loquercio, Ashish Kumar +1 authors

A fully-observable robot policy guides a partially-observable one in pursuit-evasion tasks by generating supervision, focusing on evader behavior diversity and modeling assumptions.

8fully-observablepartially-observableHF ↗arXiv ↗
22

ORES: Open-vocabulary Responsible Visual Synthesis

Minheng Ni, Chenfei Wu, Xiaodong Wang +4 authors

The Two-stage Intervention (TIN) framework uses a large-scale language model and diffusion synthesis model to enable responsible visual synthesis by avoiding forbidden visual concepts while following user queries.

8Open-vocabulary Responsible Visual SynthesisTwo-stage InterventionHF ↗arXiv ↗

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号