TensorX
返回文献探索

Paper · arXiv 2410.03027

MLP-KAN: Unifying Deep Representation and Function Learning

Yunhong He, Yifeng Xie, Zhengqing Yuan, Lichao Sun

31 upvotesOctober 3, 2024arXiv 预印本
AI 摘要

MLP-KAN integrates MLPs and KANs in a MoE architecture within a transformer framework to adaptively handle both representation and function learning tasks, achieving competitive results across diverse datasets.

Multi-Layer PerceptronsMLPsKolmogorov-Arnold NetworksKANsMixture-of-ExpertsMoEtransformer-based framework

Abstract

Recent advancements in both representation learning and function learning have demonstrated substantial promise across diverse domains of artificial intelligence. However, the effective integration of these paradigms poses a significant challenge, particularly in cases where users must manually decide whether to apply a representation learning or function learning model based on dataset characteristics. To address this issue, we introduce MLP-KAN, a unified method designed to eliminate the need for manual model selection. By integrating Multi-Layer Perceptrons (MLPs) for representation learning and Kolmogorov-Arnold Networks (KANs) for function learning within a Mixture-of-Experts (MoE) architecture, MLP-KAN dynamically adapts to the specific characteristics of the task at hand, ensuring optimal performance. Embedded within a transformer-based framework, our work achieves remarkable results on four widely-used datasets across diverse domains. Extensive experimental evaluation demonstrates its superior versatility, delivering competitive performance across both deep representation and function learning tasks. These findings highlight the potential of MLP-KAN to simplify the model selection process, offering a comprehensive, adaptable solution across various domains. Our code and weights are available at https://github.com/DLYuanGod/MLP-KAN.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
MLP-KAN: Unifying Deep Representation and Function Learning | TensorX