TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Apr 29 – May 5, 2024
本周最热123

Prometheus 2: An Open Source Language Model Specialized in Evaluating Other Language Models

Seungone Kim, Juyoung Suk, Shayne Longpre +7 authors

Prometheus 2 is an advanced open-source evaluator LM that aligns closely with human and proprietary LM judgments, supporting both direct assessment and pairwise ranking with customizable criteria.

LMsGPT-4open-source LMsevaluationHF ↗arXiv ↗

43 篇论文 · 按点赞排序

03

Octopus v4: Graph of language models

Wei Chen, Zhiyuan Li

The Octopus v4 model uses functional tokens to integrate and direct queries to task-specific open-source language models, achieving SOTA performance with models under 10B parameters.

116functional tokensOctopus v4HF ↗arXiv ↗
04

KAN: Kolmogorov-Arnold Networks

Ziming Liu, Yixuan Wang, Sachin Vaidya +5 authors

Kolmogorov-Arnold Networks (KANs) outperform Multi-Layer Perceptrons (MLPs) in accuracy and interpretability by using learnable activation functions and spline-based weights.

115Kolmogorov-Arnold NetworksKANsHF ↗arXiv ↗
05

Better & Faster Large Language Models via Multi-token Prediction

Fabian Gloeckle, Badr Youbi Idrissi, Baptiste Rozière +2 authors

Training large language models to predict multiple future tokens simultaneously enhances their sample efficiency and performance, particularly on generative benchmarks and small algorithmic tasks, with decreased inference time.

82next-token prediction lossmulti-token predictionHF ↗arXiv ↗
08

WildChat: 1M ChatGPT Interaction Logs in the Wild

Wenting Zhao, Xiang Ren, Jack Hessel +3 authors

Chatbots such as GPT-4 and ChatGPT are now serving millions of users. Despite their widespread use, there remains a lack of public datasets showcasing how these tools are used by a population of users in practice. To bridge this gap, we offered free access to ChatGPT for online users in exchange for their affirmative, consensual opt-in to anonymously collect their chat transcripts and request headers. From this, we compiled WildChat, a corpus of 1 million user-ChatGPT conversations, which consists of over 2.5 million interaction turns. We compare WildChat with other popular user-chatbot interaction datasets, and find that our dataset offers the most diverse user prompts, contains the largest number of languages, and presents the richest variety of potentially toxic use-cases for researchers to study. In addition to timestamped chat transcripts, we enrich the dataset with demographic data, including state, country, and hashed IP addresses, alongside request headers. This augmentation allows for more detailed analysis of user behaviors across different geographical regions and temporal dimensions. Finally, because it captures a broad range of use cases, we demonstrate the dataset's potential utility in fine-tuning instruction-following models. WildChat is released at https://wildchat.allen.ai under AI2 ImpACT Licenses.

66HF ↗arXiv ↗
10

Iterative Reasoning Preference Optimization

Richard Yuanzhe Pang, Weizhe Yuan, Kyunghyun Cho +3 authors

An iterative preference optimization method using a modified DPO loss improves reasoning accuracy on various datasets by optimizing winning and losing reasoning steps in Chain-of-Thought candidates.

50iterative preference optimizationChain-of-Thought (CoT)HF ↗arXiv ↗
12

Extending Llama-3's Context Ten-Fold Overnight

Peitian Zhang, Ninglu Shao, Zheng Liu +4 authors

Llama-3-8B-Instruct's context length is extended from 8K to 80K using QLoRA fine-tuning with minimal additional training samples, demonstrating significant potential for further context extension with increased computational resources.

34QLoRA fine-tuningHF ↗arXiv ↗
16

NeMo-Aligner: Scalable Toolkit for Efficient Model Alignment

Gerald Shen, Zhilin Wang, Olivier Delalleau +10 authors

NeMo-Aligner is a toolkit for aligning large language models with human values using scalable techniques like RLHF, DPO, SteerLM, and SPIN, optimized for use with hundreds of GPUs and parameter-efficient fine-tuning.

29NeMo-AlignerReinforcement Learning from Human Feedback (RLHF)HF ↗arXiv ↗
17

Self-Play Preference Optimization for Language Model Alignment

Yue Wu, Zhiqing Sun, Huizhuo Yuan +3 authors

A self-play method called SPPO for language model alignment achieves state-of-the-art performance by approximating Nash equilibrium policy in a constant-sum game setting, outperforming other approaches with limited data.

29reinforcement learning from human feedbackparametric modelsHF ↗arXiv ↗
18

AdvPrompter: Fast Adaptive Adversarial Prompting for LLMs

Anselm Paulus, Arman Zharmagambetov, Chuan Guo +2 authors

A novel method uses AdvPrompter, another LLM, to generate human-readable adversarial prompts efficiently, improving the state-of-the-art in generating unpredictable yet harmful outputs, and enhancing LLM robustness against jailbreaking attacks.

29Large Language Modelsjailbreaking attacksHF ↗arXiv ↗
21

Capabilities of Gemini Models in Medicine

Khaled Saab, Tao Tu, Wei-Hung Weng +63 authors

Med-Gemini, a family of specialized multimodal AI models, achieves state-of-the-art performance across various medical benchmarks, surpassing GPT-4 and demonstrating superior long-context and uncertainty-guided search capabilities.

25multimodaldiamond modelsHF ↗arXiv ↗
22

Visual Fact Checker: Enabling High-Fidelity Detailed Caption Generation

Yunhao Ge, Xiaohui Zeng, Jacob Samuel Huffman +3 authors

VisualFactChecker improves caption accuracy and detail through a three-step pipeline combining image-to-text models, fact-checking with large language models, and final summarization, outperforming existing methods on multiple datasets.

24image-to-text captioning modelslarge language modelHF ↗arXiv ↗
23

LEGENT: Open Platform for Embodied Agents

Zhili Cheng, Zhitong Wang, Jinyi Hu +7 authors

LEGENT is an open platform for developing embodied agents using LLMs and LMMs, featuring a rich 3D environment and a data generation pipeline that outperforms GPT-4V in embodied tasks.

22Large Language ModelsLarge Multimodal ModelsHF ↗arXiv ↗
29

SemantiCodec: An Ultra Low Bitrate Semantic Audio Codec for General Sound

Haohe Liu, Xuenan Xu, Yi Yuan +3 authors

SemantiCodec is a dual-encoder codec that compresses diverse audio types into fewer tokens using AudioMAE and k-means clustering, improved by a diffusion-model-based decoder, achieving higher semantic information and superior reconstruction quality at lower bitrates than existing codecs.

17AudioMAEk-means clusteringHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号