TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

598 篇论文 · 按点赞排序

183

SaulLM-54B & SaulLM-141B: Scaling Up Domain Adaptation for the Legal Domain

Pierre Colombo, Telmo Pires, Malik Boudiaf +7 authors

Two large legal language models, SaulLM-54B and SaulLM-141B, based on the Mixtral architecture, are introduced for domain-specific adaptation in the legal sector using continued pretraining, specialized protocols, and preference alignment with synthetic data.

66Mixtral architecturelarge language modelsHF ↗arXiv ↗
184

Understanding LLMs: A Comprehensive Overview from Training to Inference

Yiheng Liu, Hao He, Tianle Han +18 authors

The paper reviews techniques for cost-efficient training and deployment of large language models, covering aspects like data preprocessing, parallel training, model fine-tuning, and inference optimizations including model compression and memory scheduling.

66Large Language Modelspre-training tasksHF ↗arXiv ↗
186

Yi: Open Foundation Models by 01.AI

01. AI, Alex Young, Bei Chen +28 authors

The Yi model family, based on transformer architecture, showcases strong performance across benchmarks and modalities through optimized data and scalable infrastructure.

66language modelsmultimodal modelsHF ↗arXiv ↗
187

WildChat: 1M ChatGPT Interaction Logs in the Wild

Wenting Zhao, Xiang Ren, Jack Hessel +3 authors

Chatbots such as GPT-4 and ChatGPT are now serving millions of users. Despite their widespread use, there remains a lack of public datasets showcasing how these tools are used by a population of users in practice. To bridge this gap, we offered free access to ChatGPT for online users in exchange for their affirmative, consensual opt-in to anonymously collect their chat transcripts and request headers. From this, we compiled WildChat, a corpus of 1 million user-ChatGPT conversations, which consists of over 2.5 million interaction turns. We compare WildChat with other popular user-chatbot interaction datasets, and find that our dataset offers the most diverse user prompts, contains the largest number of languages, and presents the richest variety of potentially toxic use-cases for researchers to study. In addition to timestamped chat transcripts, we enrich the dataset with demographic data, including state, country, and hashed IP addresses, alongside request headers. This augmentation allows for more detailed analysis of user behaviors across different geographical regions and temporal dimensions. Finally, because it captures a broad range of use cases, we demonstrate the dataset's potential utility in fine-tuning instruction-following models. WildChat is released at https://wildchat.allen.ai under AI2 ImpACT Licenses.

65HF ↗arXiv ↗
190

Controllable Text Generation for Large Language Models: A Survey

Xun Liang, Hanyu Wang, Yezhaohui Wang +8 authors

Controllable Text Generation techniques for Large Language Models ensure predefined control conditions and high-quality text output, covering content and attribute control through various methods like retraining, fine-tuning, and latent manipulation.

65Large Language ModelsControllable Text GenerationHF ↗arXiv ↗
193

Programming Every Example: Lifting Pre-training Data Quality like Experts at Scale

Fan Zhou, Zengzhi Wang, Qian Liu +2 authors

ProX, a novel framework treating data refinement as a programming task, enhances large language model pre-training by generating fine-grained operations for each example, outperforming traditional human-crafted rule-based methods significantly across various benchmarks and domains.

64Programming Every Example (ProX)data refinementHF ↗arXiv ↗
197

Depth Anything: Unleashing the Power of Large-Scale Unlabeled Data

Lihe Yang, Bingyi Kang, Zilong Huang +3 authors

Depth Anything is a robust monocular depth estimation model built on a large-scale dataset with data augmentation and auxiliary supervision strategies, achieving state-of-the-art results on various datasets and enhancing depth-conditioned ControlNet.

64monocular depth estimationdata engineHF ↗arXiv ↗
198

Qwen2-Audio Technical Report

Yunfei Chu, Jin Xu, Qian Yang +9 authors

Qwen2-Audio, a large-scale audio-language model, enhances instruction-following and audio analysis through natural language prompts and DPO optimization.

64audio-language modelpre-training processHF ↗arXiv ↗
7 / 20

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号