TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jan 29 – Feb 4, 2024
本周最热86

OLMo: Accelerating the Science of Language Models

Dirk Groeneveld, Iz Beltagy, Pete Walsh +40 authors

OLMo, an open-sourced language model, provides comprehensive access to training data, code, and architecture, facilitating research and innovation in language modeling.

open language modellanguage modelstraining datatraining codeHF ↗arXiv ↗

50 篇论文 · 按点赞排序

07

RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval

Parth Sarthi, Salman Abdullah, Aditi Tuli +3 authors

The RAPTOR model enhances retrieval-augmented language models by recursively embedding, clustering, and summarizing text chunks, leading to better performance on question-answering tasks involving complex reasoning.

49retrieval-augmented language modelsRecursive embeddingHF ↗arXiv ↗
08

Weaver: Foundation Models for Creative Writing

Tiannan Wang, Jiamin Chen, Qingrui Jia +43 authors

Weaver, a family of specialized large language models, achieves superior writing capabilities through pre-training and fine-tuning methods, surpasses GPT-4 in various writing tasks, and supports retrieval-augmented generation and tool usage.

46large language modelspre-trainingHF ↗arXiv ↗
18

Can Large Language Models Understand Context?

Yilun Zhu, Joel Ruben Antony Moniz, Shruti Bhargava +6 authors

The benchmark evaluates LLMs' context understanding by assessing pre-trained and quantized models across four tasks and nine datasets, highlighting the performance differences between dense and fine-tuned models.

24Large Language Models (LLMs)in-context learningHF ↗arXiv ↗
20

Efficient Exploration for LLMs

Vikranth Dwaracherla, Seyed Mohammad Asghari, Botao Hao +1 authors

Efficient exploration using double Thompson sampling and epistemic neural network uncertainty estimation significantly improves large language models with reduced query numbers.

22double Thompson samplingepistemic neural networkHF ↗arXiv ↗
22

Learning Universal Predictors

Jordi Grau-Moya, Tim Genewein, Marcus Hutter +8 authors

Meta-learning can leverage Universal Turing Machine-generated data to train neural networks for universal prediction strategies.

22Meta-learningSolomonoff InductionHF ↗arXiv ↗
24

Efficient Tool Use with Chain-of-Abstraction Reasoning

Silin Gao, Jane Dwivedi-Yu, Ping Yu +7 authors

Chain-of-Abstraction (CoA) method enhances large language models' multi-step reasoning by using abstract reasoning chains to plan and efficiently invoke tools, improving accuracy and speed in tasks like mathematical reasoning and Wiki QA.

21large language modelsLLMsHF ↗arXiv ↗
28

Advances in 3D Generation: A Survey

Xiaoyu Li, Qi Zhang, Di Kang +7 authors

A survey of fundamental methodologies in 3D generation, including 3D representations, generation paradigms, datasets, applications, and challenges.

193D representationsfeedforward generationHF ↗arXiv ↗
29

Proactive Detection of Voice Cloning with Localized Watermarking

Robin San Roman, Pierre Fernandez, Alexandre Défossez +3 authors

AudioSeal is an audio watermarking technique that uses a generator/detector architecture to detect AI-generated speech with high robustness and imperceptibility, offering fast detection.

19generator/detector architecturelocalization lossHF ↗arXiv ↗
30

H2O-Danube-1.8B Technical Report

Philipp Singer, Pascal Pfeiffer, Yauhen Babakhin +4 authors

A 1.8B parameter language model trained on 1T tokens demonstrates competitive performance across benchmarks and is released publicly with supervised fine-tuning and preference optimization.

18language modelpre-trainingHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号