TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

May 1 – May 7, 2023
本周最热10

Personalize Segment Anything Model with One Shot

Renrui Zhang, Zhengkai Jiang, Ziyu Guo +5 authors

A training-free and fine-tuning variant of the Segment Anything Model (SAM), PerSAM and PerSAM-F, achieves personalized image and video segmentation using a single reference image and minimal fine-tuning, improving performance on personalized and dreambooth applications.

Segment Anything ModelSAMPerSAMPerSAM-FHF ↗arXiv ↗

18 篇论文 · 按点赞排序

04

AutoML-GPT: Automatic Machine Learning with GPT

Shujian Zhang, Chengyue Gong, Lemeng Wu +2 authors

AutoML-GPT uses LLMs to automate the entire AI training pipeline from data processing to hyperparameter tuning, achieving high performance across various domains.

4large language modelsLLMsHF ↗arXiv ↗
05

Shap-E: Generating Conditional 3D Implicit Functions

Heewoo Jun, Alex Nichol

Shap-E, a conditional generative model, generates 3D assets by training an encoder and a conditional diffusion model, achieving faster convergence and high-quality samples compared to explicit models.

4conditional generative modelimplicit functionsHF ↗arXiv ↗
12

Masked Trajectory Models for Prediction, Representation, and Control

Philipp Wu, Arjun Majumdar, Kevin Stone +4 authors

Masked Trajectory Models (MTM) are versatile sequential decision-making models that can perform tasks such as forward dynamics modeling, inverse dynamics modeling, and offline RL using generic masks and representations, matching or outperforming specialized methods.

1Masked Trajectory ModelsMTMHF ↗arXiv ↗
13

Real-Time Neural Appearance Models

Tizian Zeltner, Fabrice Rousselle, Andrea Weidlich +7 authors

Real-time rendering of complex scenes is achieved using learned hierarchical textures interpreted by neural decoders, supported by hardware-accelerated tensor operations in ray tracing shaders.

1hierarchical texturesneural decodersHF ↗arXiv ↗
14

Tracking through Containers and Occluders in the Wild

Basile Van Hoorick, Pavel Tokmakov, Simon Stent +2 authors

TCOW is a benchmark and model for visual tracking in cluttered environments with heavy occlusions and containment, using a mixture of synthetic and real datasets to evaluate transformer-based video models' performance in understanding object permanence.

1transformer-based video modelsobject permanenceHF ↗arXiv ↗

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号