TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Dec 11 – Dec 17, 2023

50 篇论文 · 按点赞排序

33

Vision-Language Models as a Source of Rewards

Kate Baumli, Satinder Baveja, Feryal Behbahani +23 authors

Researchers explore using off-the-shelf vision-language models to derive rewards for reinforcement learning agents, demonstrating improved performance across various language-based visual goals.

12reinforcement learninggeneralist agentsHF ↗arXiv ↗
37

Steering Llama 2 via Contrastive Activation Addition

Nina Rimsky, Nick Gabrieli, Julian Schulz +3 authors

Contrastive Activation Addition (CAA) modifies model activations to steer language model behavior with high precision and insight into high-level concept representation.

12Contrastive Activation AdditionCAAHF ↗arXiv ↗
39

COLMAP-Free 3D Gaussian Splatting

Yang Fu, Sifei Liu, Amey Kulkarni +3 authors

The paper presents a method for novel view synthesis and camera pose estimation using 3D Gaussian Splatting, eliminating the need for pre-processed camera poses by processing input frames sequentially.

11Neural Radiance Fields (NeRFs)3D Gaussian SplattingHF ↗arXiv ↗
42

Customizing Motion in Text-to-Video Diffusion Models

Joanna Materzynska, Josef Sivic, Eli Shechtman +3 authors

The method fine-tunes text-to-video models to incorporate customized motions from a few video samples, supports multi-person and multimodal customization, and includes a quantitative evaluation approach.

11text-to-video modelsfinetuningHF ↗arXiv ↗
44

Interfacing Foundation Models' Embeddings

Xueyan Zou, Linjie Li, Jianfeng Wang +9 authors

FIND is a transformer interface that aligns foundation model embeddings for unified image segmentation and dataset-level retrieval without tuning the foundation models.

10transformerfoundation modelsHF ↗arXiv ↗
45

MVDD: Multi-View Depth Diffusion Models

Zhen Wang, Qiangeng Xu, Feitong Tan +6 authors

MVDD is a denoising diffusion model using multi-view depth representations to generate high-quality 3D shape point clouds and meshes, achieving state-of-the-art results in 3D shape generation and depth completion.

10denoising diffusion modelsmulti-view depthHF ↗arXiv ↗
46

PathFinder: Guided Search over Multi-Step Reasoning Paths

Olga Golovneva, Sean O'Brien, Ramakanth Pasunuru +4 authors

PathFinder, a tree-search-based reasoning model, improves performance on multi-step reasoning tasks by integrating dynamic decoding, constraints, pruning, and ranking.

10chain-of-thought promptinglarge language modelsHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号