TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jun 26 – Jul 2, 2023
本周最热97

Long-range Language Modeling with Self-retrieval

Ohad Rubin, Jonathan Berant

The Retrieval-Pretrained Transformer jointly trains a retrieval-augmented language model to model long texts by fusing retrieved information for improved prediction.

retrieval-augmented language modelsRetrieval-Pretrained Transformerquery representationssemantic objectiveHF ↗arXiv ↗

50 篇论文 · 按点赞排序

05

On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes

Rishabh Agarwal, Nino Vieillard, Yongchao Zhou +4 authors

Generalized Knowledge Distillation addresses distribution mismatch in auto-regressive sequence models by training the student on its own output sequences and integrates with RL fine-tuning for effective distillation in tasks like summarization and instruction-tuning.

38Knowledge distillationauto-regressive sequence modelsHF ↗arXiv ↗
09

MotionGPT: Human Motion as a Foreign Language

Biao Jiang, Xin Chen, Wen Liu +3 authors

MotionGPT, a unified motion-language model using discrete vector quantization, achieves top performance in various motion-related tasks by treating motion as a language similar to text.

28discrete vector quantizationmotion tokensHF ↗arXiv ↗
10

Generate Anything Anywhere in Any Scene

Yuheng Li, Haotian Liu, Yangming Wen +1 authors

A text-to-image diffusion model is enhanced with adapter layers and regionally-guided sampling to achieve controlled generation of personalized objects with high fidelity.

23diffusion modelsentanglement issuesHF ↗arXiv ↗
15

Scaling MLPs: A Tale of Inductive Bias

Gregor Bachmann, Sotiris Anagnostidis, Thomas Hofmann

MLPs achieve competitive performance on vision tasks with large-scale pre-training, challenging the narrative that inductive bias is necessary for high accuracy.

17multi-layer perceptron (MLP)inductive biasHF ↗arXiv ↗
23

Language models are weak learners

Hariharan Manikandan, Yiding Jiang, J Zico Kolter

Prompt-based large language models can effectively serve as weak learners in boosting algorithms, enhancing performance on tabular data with limited samples compared to traditional methods.

11weak learnerboostingHF ↗arXiv ↗
24

System-Level Natural Language Feedback

Weizhe Yuan, Kyunghyun Cho, Jason Weston

A framework is proposed to utilize system-level natural language feedback to improve model design and performance through metric design and prompt refinement, with case studies showing its effectiveness in search and dialog generation.

11human-in-the-loopmetric designHF ↗arXiv ↗
25

OpenMask3D: Open-Vocabulary 3D Instance Segmentation

Ayça Takmaz, Elisabetta Fedele, Robert W. Sumner +3 authors

OpenMask3D is a zero-shot approach for open-vocabulary 3D instance segmentation using class-agnostic instance masks and multi-view fusion of CLIP-based image embeddings.

11open-vocabulary 3D instance segmentation3D instance masksHF ↗arXiv ↗
28

Supervised Pretraining Can Learn In-Context Reinforcement Learning

Jonathan N. Lee, Annie Xie, Aldo Pacchiano +4 authors

Transformers pretrained with Decision-Pretrained Transformer (DPT) exhibit strong in-context learning capabilities in reinforcement learning, achieving exploration, generalization, and faster learning compared to algorithms used for pretraining.

9transformersreinforcement learningHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号