Addition is All You Need for Energy-efficient Language Models
Hongyin Luo, Wei Sun
A new integer addition-based algorithm reduces energy consumption in floating point tensor multiplications with comparable precision across various tasks.
Explore · 每周精选
发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。
Tianzhu Ye, Li Dong, Yuqing Xia +4 authors
Diff Transformer improves large language models by selectively focusing attention on relevant context and reducing noise, leading to better performance in scaling, long-context modeling, key information retrieval, and in-context learning.
50 篇论文 · 按点赞排序
Hongyin Luo, Wei Sun
A new integer addition-based algorithm reduces energy consumption in floating point tensor multiplications with comparable precision across various tasks.
Dongxu Li, Yudong Liu, Haoning Wu +7 authors
Aria is an open multimodal native AI model with best-in-class performance across various tasks, designed with a mixture-of-experts architecture and pre-trained through a four-stage pipeline.
Eilam Shapira, Omer Madmon, Itamar Reinman +3 authors
A benchmark for evaluating Large Language Models in strategic, language-based games provides insights into their rationality, mimicry of human behavior, and impact of economic characteristics on performance and outcome.
Renjie Pi, Jianshu Zhang, Tianyang Han +3 authors
A new framework called Personalized Visual Instruction Tuning (PVIT) enhances multimodal large language models to recognize and engage with specific individuals in images, utilizing a curated dataset and benchmarks for evaluation.
Pravesh Agrawal, Szymon Antoniak, Emma Bou Hanna +34 authors
Pixtral-12B, a 12-billion-parameter multimodal language model, excels in both natural language and image understanding, surpassing larger models and introducing an open-source benchmark for evaluation.
Di Zhang, Jianbo Wu, Jingdi Lei +9 authors
LLaMA-Berry enhances LLMs for mathematical reasoning by combining Monte Carlo Tree Search with Self-Refine and leveraging a Pairwise Preference Reward Model to optimize search efficiency and problem-solving capability.
Siyu Zhou, Tianyi Zhou, Yijun Yang +4 authors
A neurosymbolic approach aligns large language models with their environment through rule learning, improving model-based agent performance in open-world tasks.
Zimu Lu, Aojun Zhou, Ke Wang +5 authors
A novel method for generating mathematical code with reasoning steps enhances mathematical reasoning abilities in large language models using a comprehensive dataset named MathCode-Pile.
Yushen Chen, Zhikang Niu, Ziyang Ma +5 authors
F5-TTS, a fully non-autoregressive text-to-speech system, improves E2 TTS by leveraging ConvNeXt and Sway Sampling for better performance and efficiency.
Hadas Orgad, Michael Toker, Zorik Gekhman +4 authors
LLMs internally encode detailed truthfulness information about their outputs, which can be used to detect and predict errors but shows variability across datasets and does not always align with the final output.
Fanqing Meng, Jiaqi Liao, Xinyu Tan +7 authors
PhyGenBench and PhyGenEval assess text-to-video models' understanding of physical commonsense, revealing shortcomings that cannot be fully addressed by scaling models or prompt engineering.
Xinchen Zhang, Ling Yang, Guohao Li +6 authors
IterComp, a novel framework aggregating preferences from multiple diffusion models using iterative feedback learning, significantly improves compositional text-to-image generation across various metrics.
Yang Jin, Zhicheng Sun, Ningyuan Li +8 authors
A unified pyramidal flow matching algorithm with a single Diffusion Transformer enables efficient high-quality video generation by interlinking pyramid stages and compressing full-resolution history.
Qidong Huang, Xiaoyi Dong, Pan Zhang +6 authors
MIR, a new metric for evaluating multi-modal pre-training in Large Vision Language Models, effectively correlates with benchmark performance, is robust across data, and generalizes well, aiding in data selection, training strategy, and architecture design.
Junpeng Yue, Xinru Xu, Börje F. Karlsson +1 authors
The proposed MART method enhances embodied agents by fine-tuning an MLLM retriever with preference learning to prioritize task-effective trajectories, improving success rates in unseen environments.
Siyuan Li, Juanxi Tian, Zedong Wang +6 authors
The interplay between vision backbones and optimizers, termed Backbone-Optimizer Coupling Bias, significantly impacts pre-training and fine-tuning of vision models, with different architectures aligning with specific optimizer types.
Jingwei Zuo, Maksim Velikanov, Dhia Eddine Rhaiem +4 authors
Falcon Mamba 7B, a pure Mamba architecture language model, outperforms leading Transformer and hybrid models with faster inference and lower memory usage.
Mengzhao Chen, Yi Liu, Jiahao Wang +3 authors
PrefixQuant improves quantization efficiency in LLMs by isolating outlier tokens, enabling static quantization to outperform dynamic quantization in both speed and accuracy.
Yunhong He, Yifeng Xie, Zhengqing Yuan +1 authors
MLP-KAN integrates MLPs and KANs in a MoE architecture within a transformer framework to adaptively handle both representation and function learning tasks, achieving competitive results across diverse datasets.
Shuofei Qiao, Runnan Fang, Zhisong Qiu +6 authors
WorFBench and WorFEval create a comprehensive framework for evaluating large language models' workflow generation, uncovering gaps in sequence and graph planning and demonstrating improved performance in downstream tasks.
Dohun Lee, Bryan S Kim, Geon Yeong Park +1 authors
VideoGuide enhances temporal consistency and image fidelity in text-to-video generation by leveraging pretrained video diffusion models during the inference process without additional training.
Rose E. Wang, Ana T. Ribeiro, Carly D. Robinson +2 authors
Tutor CoPilot, a Human-AI system, enhances tutoring effectiveness by providing expert-like guidance, resulting in better student mastery and improved pedagogical strategies.
Yihong Dong, Ge Li, Yongding Tao +6 authors
A new Fourier-based network architecture, FAN, efficiently models periodic phenomena with fewer parameters and demonstrates superior performance across various tasks.
Saaket Agashe, Jiuzhou Han, Shuyu Gan +3 authors
Agent S, a framework for autonomous GUI interactions, enhances task automation with experience-augmented hierarchical planning and Multimodal Large Language Models.
Jiatao Gu, Yuyang Wang, Yizhe Zhang +5 authors
DART, a transformer-based model combining autoregressive and diffusion components in a non-Markovian framework, offers competitive performance on image and text-to-image generation tasks.
Yaniv Leviathan, Matan Kalman, Yossi Matias
Selective Attention reduces unnecessary context elements in the attention mechanism, improving language modeling performance and decreasing memory and compute requirements.
Hanrong Ye, Haotian Zhang, Erik Daxberger +9 authors
A multimodal foundation model for egocentric video understanding is developed with a large QA dataset, a challenging benchmark, and a specialized architecture featuring a "Memory Pointer Prompting" mechanism.
Xiang Liu, Peijie Dong, Xuming Hu +1 authors
A new benchmark, LongGenBench, evaluates the long-context generation capabilities of large language models, revealing varying performance degradation across different models and model series.
Iman Mirzadeh, Keivan Alizadeh, Hooman Shahrokhi +3 authors
LSSMs exhibit variability and fragility in mathematical reasoning, as evidenced by performance declines on GSM-Symbolic benchmark questions with altered numerical values and additional clauses.
北京市昌平区探索星信息技术及软件开发工作室
京ICP备2026059466号