TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

598 篇论文 · 按点赞排序

422

Kolmogorov-Arnold Transformer

Xingyi Yang, Xinchao Wang

The Kolmogorov-Arnold Transformer replaces MLP layers with Kolmogorov-Arnold Network layers to enhance transformers, overcoming challenges related to inference speed, computation efficiency, and weight initialization through rational basis, group learning, and variance-preserving techniques.

45Kolmogorov-Arnold TransformerKATHF ↗arXiv ↗
426

Branch-Train-MiX: Mixing Expert LLMs into a Mixture-of-Experts LLM

Sainbayar Sukhbaatar, Olga Golovneva, Vasu Sharma +8 authors

Branch-Train-MiX (BTX) method enhances Large Language Models by asynchronously training experts in parallel and integrating them using Mixture-of-Expert layers with token-level routing for improved accuracy and efficiency.

45Large Language ModelsBranch-Train-MiXHF ↗arXiv ↗
433

ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Xiwei Hu, Rui Wang, Yixiao Fang +3 authors

ELLA, an Efficient Large Language Model Adapter, enhances text-to-image diffusion models by integrating powerful Large Language Models through a Timestep-Aware Semantic Connector, improving dense prompt comprehension and generation quality.

45diffusion modelstext-to-image generationHF ↗arXiv ↗
434

EVLM: An Efficient Vision-Language Model for Visual Understanding

Kaibing Chen, Dong Shen, Hanwen Zhong +14 authors

A multi-modal language model using cross-attention, hierarchical ViT features, and Mixture of Experts mechanism achieves competitive performance in image and video captioning tasks with reduced computational costs.

45cross-attentionhierarchical ViT featuresHF ↗arXiv ↗
435

Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages

Xiang Yue, Yueqi Song, Akari Asai +7 authors

Pangea, a multilingual multimodal LLM, is trained on a diverse 6M instruction dataset across 39 languages and evaluated using PangeaBench, demonstrating superior performance in multilingual and cross-cultural settings.

44multimodal large language models (MLLMs)PangeaHF ↗arXiv ↗
437

Foundation Models for Music: A Survey

Yinghao Ma, Anders Øland, Anton Ragni +40 authors

A review of foundation models in music, including large language models and latent diffusion models, highlights their impact, architectural choices, and the need for ethical considerations in music applications.

44large language modelslatent diffusion modelsHF ↗arXiv ↗
438

Very Large-Scale Multi-Agent Simulation in AgentScope

Xuchen Pan, Dawei Gao, Yuexiang Xie +5 authors

Enhancements to the AgentScope platform improve scalability, efficiency, and ease of use for large-scale multi-agent simulations through distributed mechanisms, flexible environments, and user-friendly tools.

44actor-based distributed mechanismmulti-agent platformHF ↗arXiv ↗
439

SpectroMotion: Dynamic 3D Reconstruction of Specular Scenes

Cheng-De Fan, Chen-Wei Chang, Yi-Ruei Liu +4 authors

SpectroMotion improves upon 3D Gaussian Splatting by integrating physical-based rendering and deformation fields, incorporating a residual correction technique and deformable environment map to accurately reconstruct and synthesize dynamic specular scenes.

443D Gaussian Splatting (3DGS)physically-based rendering (PBR)HF ↗arXiv ↗
441

Transformers meet Neural Algorithmic Reasoners

Wilfried Bounsi, Borja Ibarz, Andrew Dudzik +5 authors

A novel TransNAR model combines Transformer-based language understanding with graph neural network solvers to enhance algorithmic reasoning.

44Transformersnatural language understandingHF ↗arXiv ↗
445

KAN or MLP: A Fairer Comparison

Runpeng Yu, Weihao Yu, Xinchao Wang

A comprehensive comparison of KAN and MLP models across diverse tasks reveals that MLP generally outperforms KAN except in symbolic formula representation where KAN's B-spline activation function provides advantage, and KAN exhibits more severe forgetting issues in class-incremental continual learning.

43KANMLPHF ↗arXiv ↗
446

MindSearch: Mimicking Human Minds Elicits Deep AI Searcher

Zehui Chen, Kuikun Liu, Qiuchen Wang +4 authors

MindSearch, an LLM-based multi-agent framework, improves web information seeking and integration through parallel processing and hierarchical retrieval, achieving better performance than existing solutions.

43Large Language ModelsLLMsHF ↗arXiv ↗
447

Depth Pro: Sharp Monocular Metric Depth in Less Than a Second

Aleksei Bochkovskii, Amaël Delaunoy, Hugo Germain +4 authors

Depth Pro is a fast and accurate zero-shot monocular depth estimation model that generates high-resolution depth maps using a multi-scale transformer and a combined real-synthetic training protocol.

43zero-shotmonocular depth estimationHF ↗arXiv ↗
450

How to Train Data-Efficient LLMs

Noveen Sachdeva, Benjamin Coleman, Wang-Cheng Kang +6 authors

Data-efficient methods like Ask-LLM and Density sampling improve model quality and training efficiency in large language models by optimizing data selection and coverage.

43large language modelsdata-efficientHF ↗arXiv ↗
15 / 20

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号