TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Sep 9 – Sep 15, 2024

49 篇论文 · 按点赞排序

31

PiTe: Pixel-Temporal Alignment for Large Video-Language Model

Yang Liu, Pengxiang Ding, Siteng Huang +3 authors

A new video-language model, PiTe, uses trajectory-guided pixel-temporal alignment to achieve superior performance across various multimodal video tasks by leveraging a large pre-training dataset with precise object trajectories.

14Large Language Models (LLMs)Large Visual-Language Models (LVLMs)HF ↗arXiv ↗
33

Robot Utility Models: General Policies for Zero-Shot Deployment in New Environments

Haritheja Etukuru, Norihito Naka, Zijin Hu +7 authors

Robot models, particularly those trained with large amounts of data, have recently shown a plethora of real-world manipulation and navigation capabilities. Several independent efforts have shown that given sufficient training data in an environment, robot policies can generalize to demonstrated variations in that environment. However, needing to finetune robot models to every new environment stands in stark contrast to models in language or vision that can be deployed zero-shot for open-world problems. In this work, we present Robot Utility Models (RUMs), a framework for training and deploying zero-shot robot policies that can directly generalize to new environments without any finetuning. To create RUMs efficiently, we develop new tools to quickly collect data for mobile manipulation tasks, integrate such data into a policy with multi-modal imitation learning, and deploy policies on-device on Hello Robot Stretch, a cheap commodity robot, with an external mLLM verifier for retrying. We train five such utility models for opening cabinet doors, opening drawers, picking up napkins, picking up paper bags, and reorienting fallen objects. Our system, on average, achieves 90% success rate in unseen, novel environments interacting with unseen objects. Moreover, the utility models can also succeed in different robot and camera set-ups with no further data, training, or fine-tuning. Primary among our lessons are the importance of training data over training algorithm and policy class, guidance about data scaling, necessity for diverse yet high-quality demonstrations, and a recipe for robot introspection and retrying to improve performance on individual environments. Our code, data, models, hardware designs, as well as our experiment and deployment videos are open sourced and can be found on our project website: https://robotutilitymodels.com

14multi-modal imitation learningon-device deploymentHF ↗arXiv ↗
45

UniDet3D: Multi-dataset Indoor 3D Object Detection

Maksim Kolodiazhnyi, Anna Vorontsova, Matvey Skripkin +2 authors

A 3D object detection model trained on a combined indoor dataset achieves superior performance across multiple benchmarks using a unified label space and transformer encoder architecture.

83D object detectionpoint cloudsHF ↗arXiv ↗
46

Generative Hierarchical Materials Search

Sherry Yang, Simon Batzner, Ruiqi Gao +7 authors

GenMS combines a language model, diffusion model, and graph neural network to generate crystal structures from natural language, improving satisfaction of user requests and energy efficiency.

7language-to-structure generationmulti-objective optimizationHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号