MindAgent: Emergent Gaming Interaction
Ran Gong, Qiuyuan Huang, Xiaojian Ma +8 authors
MindAgent infrastructure evaluates multi-agent coordination and collaboration using LLMs in gaming scenarios, introducing CUISINEWORLD and CoS metric.
Explore · 每周精选
发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。
50 篇论文 · 按点赞排序
Ran Gong, Qiuyuan Huang, Xiaojian Ma +8 authors
MindAgent infrastructure evaluates multi-agent coordination and collaboration using LLMs in gaming scenarios, introducing CUISINEWORLD and CoS metric.
Stéphane d'Ascoli, Samy Bengio, Josh Susskind +1 authors
Boolformer, a Transformer architecture, achieves competitive symbolic regression of Boolean functions on real-world datasets and gene regulatory networks with significant speed improvements over genetic algorithms.
Nolan Dey, Daria Soboleva, Faisal Al-Khateeb +11 authors
BTLM-3B-8K, a state-of-the-art 3 billion parameter open-source language model, outperforms existing 3B models and rivals 7B models, offering excellent long context performance and reduced memory and inference compute requirements.
Xiangru Tang, Yiming Zong, Yilun Zhao +2 authors
Structure-aware fine-tuning improves Large Language Models' ability to generate complex structured data by reducing formatting errors and enhancing adherence to natural language constraints.
Zhiqiang Shen, Tianhua Tao, Liqun Ma +5 authors
Research on SlimPajama-DC reveals the impact of various data deduplication strategies and proportions on large language model training, showing improved performance with optimal data diversity.
Baolin Peng, Linfeng Song, Ye Tian +3 authors
Two methods, Advantage Model and Selective Rehearsal, are introduced to enhance the stability and performance of RLHF training for Large Language Models.
Anurag Ajay, Seungwook Han, Yilun Du +7 authors
HiP, a compositional foundation model, integrates language, vision, and action models to plan and execute complex, long-horizon manipulation tasks through hierarchical reasoning and iterative refinement.
Luoyi Sun, Xuenan Xu, Mengyue Wu +1 authors
An automatic pipeline creates a large audio-text dataset, Auto-ACD, improving performance on audio-language retrieval, captioning, and classification tasks.
Egor Lakomkin, Chunyang Wu, Yassir Fathullah +3 authors
A novel method using LLMs for contextualizing speech recognition models achieves significant performance improvements with minimal additional parameters.
Xinhao Mei, Varun Nagaraja, Gael Le Lan +4 authors
FoleyGen, a V2A system employing a Transformer model with novel visual attention mechanisms, achieves superior performance by leveraging bidirectional neural audio codecs and various pretrained visual encoders.
Adam Rashid, Satvik Sharma, Chung Min Kim +4 authors
LERF-TOGO uses vision-language models and DINO features for zero-shot task-oriented grasping by extracting object masks and ranking grasps over specific parts.
Huayang Li, Siheng Li, Deng Cai +5 authors
TextBind is an annotation-free framework that enhances large language models for multimodal instruction following by generating conversations from image-caption pairs.
Rajarshi Bhowmik, Marco Ponza, Atharva Tendle +5 authors
Fine-tuning medium-sized language models with a cross-encoder style significantly improves salient entity detection compared to feature engineering approaches.
Hao-Jun Michael Shi, Tsung-Hsien Lee, Shintaro Iwasaki +5 authors
Shampoo, an AdaGrad-based method using block-diagonal preconditioners with Kronecker product approximations, enhances neural network training performance in PyTorch, especially distributed multi-GPU environments.
Yi Yuan, Haohe Liu, Xubo Liu +3 authors
A retrieval-augmented approach enhances text-to-audio generation by addressing class imbalance and improving generation performance for rare audio classes.
Jack Lanchantin, Sainbayar Sukhbaatar, Gabriel Synnaeve +3 authors
A new data generator for machine reasoning is introduced, integrating with an embodied agent, and tests baseline models including pre-trained language models and graph-structured Transformers on a knowledge-graph representation of the database.
Nuri Ryu, Minsu Gong, Geonung Kim +2 authors
POP3D is a framework that creates a 360-degree 3D model from a single image using a combination of depth prediction, space carving, generative modeling, and neural implicit surfaces.
Aleksandar Stanić, Dylan Ashley, Oleg Serikov +5 authors
A protocol for comparing language models based on equivalent computational resources evaluates GPT-2 and a new high-throughput LSTM, showing predictable scaling and intersection in performance at 50,000 accelerator hours.
Gael Le Lan, Varun Nagaraja, Ernie Chang +5 authors
A new stack-and-delay decoding strategy for hierarchical token stacks in music generation improves inference speed without significantly compromising quality.
Sarkar Snigdha Sarathi Das, Chirag Shah, Mengting Wan +5 authors
S3-DST, a structured prompting technique with Pre-Analytical Recollection, enhances joint dialogue segmentation and state tracking in open-domain dialogs handled by LLM-based systems.
北京市昌平区探索星信息技术及软件开发工作室
京ICP备2026059466号