Multi-Agent System for Comprehensive Soccer Understanding
Jiayuan Rao, Zifeng Li, Haoning Wu +3 authors
A framework for comprehensive soccer understanding includes a knowledge base, benchmark, and multi-agent system for reasoning and evaluation.
Explore · 每周精选
发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。
50 篇论文 · 按点赞排序
Jiayuan Rao, Zifeng Li, Haoning Wu +3 authors
A framework for comprehensive soccer understanding includes a knowledge base, benchmark, and multi-agent system for reasoning and evaluation.
Jiarui Yao, Yifan Hao, Hanning Zhang +4 authors
GVM-RAFT, a dynamic sampling strategy for chain-of-thought reasoning in large language models, improves convergence and accuracy by adaptively allocating computational resources.
Xingyu Zheng, Yuye Li, Haoran Chu +7 authors
This study evaluates the impact of low-bit quantization on Qwen3, a state-of-the-art LLM, across various bit-widths and datasets, revealing performance trade-offs and suggesting areas for further research to improve quantization methods.
Qingkai Fang, Yan Zhou, Shoutao Guo +2 authors
LLaMA-Omni 2, a series of speech language models with parameters ranging from 0.5B to 14B, achieves high-quality real-time speech interaction through a speech encoder and autoregressive streaming speech decoder, outperforming models like GLM-4-Voice with significantly less training data.
Beichen Wen, Haozhe Xie, Zhaoxi Chen +2 authors
This survey examines the state-of-the-art methods in 3D scene generation, categorizing them into procedural, neural 3D, image-based, and video-based paradigms, and discusses current challenges and future directions.
Kai Ruan, Mowen Huang, Ji-Rong Wen +1 authors
New benchmark evaluates LLMs in decentralized coordination tasks under limited information, highlighting challenges and potential for future systems.
Dmitriy Shopkhoev, Ammar Ali, Magauiya Zhussip +4 authors
ReplaceMe is a training-free depth pruning method that replaces transformer blocks with linear operations using calibration data, achieving high compression ratios with minimal performance loss.
Runyi Yu, Yinhuai Wang, Qihan Zhao +4 authors
The approach uses data augmentation techniques and adaptive sampling to improve robust skill acquisition and generalization in Reinforcement Learning from Interaction Demonstration despite noisy and sparse demonstrations.
Minzheng Wang, Yongbin Li, Haobo Wang +6 authors
A novel adaptive mode learning framework significantly enhances social intelligence simulation by dynamically selecting reasoning modes based on context, reducing token usage and improving performance compared to existing methods.
Cfir Avraham Hadar, Omer Shubi, Yoav Meiri +1 authors
LLMs can ascertain readers' specific information goals through eye movement data, as demonstrated by successful goal classification and reconstruction tasks.
Chunyu Xie, Bin Wang, Fanjing Kong +5 authors
FG-CLIP enhances fine-grained understanding in multimodal tasks by leveraging large multimodal models, a high-quality dataset with detailed captions, and hard fine-grained negative samples.
Haiyang Zhou, Wangbo Yu, Jiawen Guan +3 authors
HoloTime uses a two-stage panoramic diffusion model and a space-time depth estimation method to generate high-fidelity 4D assets for immersive VR and AR experiences.
Qianchu Liu, Sheng Zhang, Guanghui Qin +9 authors
X-Reasoner, a vision-language model post-trained on general-domain text, achieves strong reasoning capabilities across modalities and domains, with X-Reasoner-Med outperforming existing models on medical benchmarks.
John Yang, Kilian Leret, Carlos E. Jimenez +7 authors
SWE-smith automates the generation of large-scale software engineering training data and improves the performance of language models on automated software engineering tasks.
Ming Li, Xin Gu, Fan Chen +4 authors
The paper presents a novel solution for improving instruction-based image editing by rectifying and enhancing editing instructions through contrastive supervision, thereby outperforming existing methods on benchmarks.
Biao Gong, Cheng Zou, Dandan Zheng +13 authors
Ming-Lite-Uni, an open-source multimodal framework, integrates vision and language using unified visual generators and autoregressive models, demonstrating strong performance in text-to-image generation and image editing.
Ahmed Abdelreheem, Filippo Aleotti, Jamie Watson +6 authors
A new task for placing 3D assets in real 3D scenes based on textual prompts is introduced, with a benchmark and dataset for evaluating 3D LLMs.
Haibo Wang, Bo Feng, Zhengfeng Lai +6 authors
StreamBridge is a framework that enhances offline Video-LLMs for streaming capabilities through a memory buffer with compression and a proactive response model, demonstrating superior performance on video understanding tasks.
Stef De Sabbata, Stefano Mizzaro, Kevin Roitero
A framework for understanding how Large Language Models process geographical information using spatial analysis and mechanistic interpretability techniques.
Ilan Strauss, Isobel Moure, Tim O'Reilly +1 authors
Research by leading AI organizations focuses more on pre-deployment stages like model alignment and testing over deployment issues such as bias, with significant gaps existing in high-risk areas like healthcare and finance.
北京市昌平区探索星信息技术及软件开发工作室
京ICP备2026059466号