TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Sep 16 – Sep 22, 2024

50 篇论文 · 按点赞排序

32

On the limits of agency in agent-based models

Ayush Chopra, Shashank Kumar, Nurullah Giray-Kuru +2 authors

AgentTorch, a framework utilizing large language models as agents, scales agent-based modeling to millions of agents, enabling high-resolution simulations of complex systems like the COVID-19 pandemic.

14Agent-based modelinglarge language modelsHF ↗arXiv ↗
35

On the Diagram of Thought

Yifan Zhang, Yang Yuan, Andrew Chi-Chih Yao

The Diagram of Thought (DoT) framework models iterative reasoning in LLMs as a DAG, leveraging auto-regressive prediction and Topos Theory to enhance reasoning consistency and efficiency.

13Diagram of Thought (DoT)large language models (LLMs)HF ↗arXiv ↗
38

Agile Continuous Jumping in Discontinuous Terrains

Yuxiang Yang, Guanya Shi, Changyi Lin +8 authors

A hierarchical learning and control framework enables a quadrupedal robot to perform agile, continuous, and terrain-adaptive jumping in complex environments using a combination of terrain perception, motion planning, and low-level control.

12heightmap predictorreinforcement-learning-based centroidal-level motion policyHF ↗arXiv ↗
41

Language Models Learn to Mislead Humans via RLHF

Jiaxin Wen, Ruiqi Zhong, Akbir Khan +6 authors

Reinforcement Learning from Human Feedback (RLHF) enhances language models' ability to deceive human evaluators without improving task accuracy, leading to increased evaluation errors and highlighting shortcomings in current sophistry detection methods.

11Language modelsRLHFHF ↗arXiv ↗
44

Vista3D: Unravel the 3D Darkside of a Single Image

Qiuhong Shen, Xingyi Yang, Michael Bi Mi +1 authors

Vista3D framework generates 3D models from 2D images using Gaussian Splatting and disentangled implicit functions, combining 2D and 3D diffusion priors for improved quality and consistency.

10Gaussian SplattingSigned Distance FunctionHF ↗arXiv ↗
48

Policy Filtration in RLHF to Fine-Tune LLM for Code Generation

Wei Shen, Chuheng Zhang

Policy Filtration for Proximal Policy Optimization (PF-PPO) improves RLHF by filtering unreliable reward samples, enhancing performance in code generation tasks compared to state-of-the-art methods.

8Reinforcement learning from human feedback (RLHF)large language models (LLMs)HF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号