TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jul 3 – Jul 9, 2023

44 篇论文 · 按点赞排序

37

Preference Ranking Optimization for Human Alignment

Feifan Song, Bowen Yu, Minghao Li +4 authors

A new method, Preference Ranking Optimization (PRO), is introduced to align LLMs with human preferences by extending the Bradley-Terry model to longer preference rankings, outperforming existing reinforcement learning techniques.

6reinforcement learning from human feedback (RLHF)reward modelHF ↗arXiv ↗
39

Elastic Decision Transformer

Yueh-Hua Wu, Xiaolong Wang, Masashi Hamaya

Elastic Decision Transformer improves upon Decision Transformer by facilitating trajectory stitching through dynamic history adjustment, achieving performance on par with Q Learning in multi-task scenarios.

5Elastic Decision TransformerDecision TransformerHF ↗arXiv ↗
41

Embodied Task Planning with Large Language Models

Zhenyu Wu, Ziwei Wang, Xiuwei Xu +2 authors

TaPA integrates LLMs with visual perception to generate feasible and effective plans for embodied agents in complex environments, outperforming other models like LLaVA and GPT-3.5.

5TAsk Planing Agent (TaPA)grounded planningHF ↗arXiv ↗
44

ReMaX: Relaxing for Better Training on Efficient Panoptic Segmentation

Shuyang Sun, Weijun Wang, Qihang Yu +3 authors

ReMaX introduces relaxation techniques to improve mask transformer training for panoptic segmentation without extra inference costs, achieving state-of-the-art results with efficient backbones on benchmarks like COCO, ADE20K, and Cityscapes.

3mask transformerspanoptic segmentationHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号