TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jun 19 – Jun 25, 2023
本周最热159

Textbooks Are All You Need

Suriya Gunasekar, Yi Zhang, Jyoti Aneja +16 authors

A new compact Transformer-based large language model for code, phi-1, achieves high accuracy on coding benchmarks despite having fewer parameters than competing models.

Transformer-basedHumanEvalMBPPparameter-efficient fine-tuningHF ↗arXiv ↗

37 篇论文 · 按点赞排序

04

Fast Segment Anything

Xu Zhao, Wenchao Ding, Yongqi An +5 authors

A speed-up method using a CNN detector with an instance segmentation branch achieves comparable performance to SAM with 50 times faster runtime by training on a small subset of SAM's dataset.

36segment anything modelSAMHF ↗arXiv ↗
07

Training Transformers with 4-bit Integers

Haocheng Xi, Changhao Li, Jianfei Chen +1 authors

A novel method for training transformers with 4-bit quantization achieves competitive accuracy and accelerates training on current GPUs.

23activation quantizationweight quantizationHF ↗arXiv ↗
08

Demystifying GPT Self-Repair for Code Generation

Theo X. Olausson, Jeevana Priya Inala, Chenglong Wang +2 authors

GPT-4 demonstrates superior self-repair capabilities on APPS dataset compared to GPT-3.5, with performance boosted when feedback is provided by GPT-4 or human programmers.

21Large Language ModelsLLMsHF ↗arXiv ↗
10

HomeRobot: Open-Vocabulary Mobile Manipulation

Sriram Yenamandra, Arun Ramachandran, Karmesh Yadav +15 authors

The HomeRobot OVMM benchmark evaluates robots' ability to manipulate unseen objects in new environments using a combination of simulated and real-world testing.

17Open-Vocabulary Mobile ManipulationpercpetionHF ↗arXiv ↗
11

Scaling Open-Vocabulary Object Detection

Matthias Minderer, Alexey Gritsenko, Neil Houlsby

OWL-ST self-training enhances OWLv2 model's open-vocabulary object detection performance by scaling up to over 1B examples, significantly improving rare class detection.

16self-trainingpseudo-box annotationsHF ↗arXiv ↗
15

Robot Learning with Sensorimotor Pre-training

Ilija Radosavovic, Baifeng Shi, Letian Fu +3 authors

A self-supervised sensorimotor pre-training method using a Transformer model on sequences of sensorimotor tokens improves robotic performance in tasks like block stacking.

14Transformersensorimotor tokensHF ↗arXiv ↗
18

DecodingTrust: A Comprehensive Assessment of Trustworthiness in GPT Models

Boxin Wang, Weixin Chen, Hengzhi Pei +16 authors

A comprehensive trustworthiness evaluation of GPT models, focusing on vulnerabilities related to toxicity, bias, adversarial robustness, privacy, ethics, and fairness, reveals that GPT-4, while generally more trustworthy, can be misled by misleading instructions.

13trustedworthiness evaluationlarge language modelsHF ↗arXiv ↗
20

Block-State Transformer

Mahan Fathi, Jonathan Pilault, Pierre-Luc Bacon +3 authors

A hybrid model combining state space models and block-wise attention outperforms Transformer-based architectures in language modeling, offering faster processing speeds.

10state space modelsBlock-State TransformerHF ↗arXiv ↗
21

Inverse Scaling: When Bigger Isn't Better

Ian R. McKenzie, Alexander Lyzhov, Michael Pieler +24 authors

Empirical evidence suggests that large language models may show inverse scaling in performance on certain tasks, leading to the importance of careful consideration of training data and objectives.

10large language modelsinverse scalingHF ↗arXiv ↗
25

GLIMMER: generalized late-interaction memory reranker

Michiel de Jong, Yury Zemlyanskiy, Nicholas FitzGerald +3 authors

GLIMMER enhances memory-augmented language models by using a shallow reranker and multi-task training, improving retrieval quality and speed over LUMEN and FiD on knowledge-intensive tasks.

9memory-augmentationlanguage modelsHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号