TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Dec 16 – Dec 22, 2024

50 篇论文 · 按点赞排序

37

Wonderland: Navigating 3D Scenes from a Single Image

Hanwen Liang, Junli Cao, Vidit Goel +6 authors

A novel pipeline using latents from a video diffusion model to predict 3D scenes from single images achieves high quality and efficiency, outperforming existing methods.

16video diffusion model3D Gaussian SplattingsHF ↗arXiv ↗
42

Affordance-Aware Object Insertion via Mask-Aware Dual Diffusion

Jixuan He, Wanhua Li, Ye Liu +3 authors

The Mask-Aware Dual Diffusion (MADD) model enables seamless object insertion into scenes by explicitly modeling the insertion mask in the diffusion process, addressing data limitations and generalizing well to real-world images.

14Affordanceaffordance-aware object insertionHF ↗arXiv ↗
43

AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling

Zihan Liu, Yang Chen, Mohammad Shoeybi +2 authors

In this paper, we introduce AceMath, a suite of frontier math models that excel in solving complex math problems, along with highly effective reward models capable of evaluating generated solutions and reliably identifying the correct ones. To develop the instruction-tuned math models, we propose a supervised fine-tuning (SFT) process that first achieves competitive performance across general domains, followed by targeted fine-tuning for the math domain using a carefully curated set of prompts and synthetically generated responses. The resulting model, AceMath-72B-Instruct greatly outperforms Qwen2.5-Math-72B-Instruct, GPT-4o and Claude-3.5 Sonnet. To develop math-specialized reward model, we first construct AceMath-RewardBench, a comprehensive and robust benchmark for evaluating math reward models across diverse problems and difficulty levels. After that, we present a systematic approach to build our math reward models. The resulting model, AceMath-72B-RM, consistently outperforms state-of-the-art reward models. Furthermore, when combining AceMath-72B-Instruct with AceMath-72B-RM, we achieve the highest average rm@8 score across the math reasoning benchmarks. We will release model weights, training data, and evaluation benchmarks at: https://research.nvidia.com/labs/adlr/acemath

13supervised fine-tuning (SFT)instruction-tunedHF ↗arXiv ↗
50

AnySat: An Earth Observation Model for Any Resolutions, Scales, and Modalities

Guillaume Astruc, Nicolas Gonthier, Clement Mallet +1 authors

Anysat uses joint embedding predictive architecture and resolution-adaptive spatial encoders to train a unified multimodal model on diverse Earth observation datasets, achieving state-of-the-art performance across various environment monitoring tasks.

11joint embedding predictive architectureJEPAHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号