TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

598 篇论文 · 按点赞排序

451

TransformerFAM: Feedback attention is working memory

Dongseong Hwang, Weiran Wang, Zhuoyuan Huo +2 authors

Feedback Attention Memory (FAM) enhances Transformer architecture by enabling long-context processing without additional weights, significantly improving performance on large sequences across various model sizes.

43TransformersFeedback Attention MemoryHF ↗arXiv ↗
471

McEval: Massively Multilingual Code Evaluation

Linzheng Chai, Shukai Liu, Jian Yang +15 authors

A multilingual code benchmark covering 40 programming languages with 16K test samples is introduced to advance code language model research, along with a multilingual coder model and instruction corpora.

41large language modelscode understandingHF ↗arXiv ↗
472

Bootstrapping Language Models with DPO Implicit Rewards

Changyu Chen, Zichen Liu, Chao Du +5 authors

A novel method using the implicit reward model from Direct Preference Optimization (DPO) to iteratively improve the alignment of large language models, achieving superior performance without external feedback.

41direct preference optimization (DPO)reinforcement learning from human feedback (RLHF)HF ↗arXiv ↗
473

Block Transformer: Global-to-Local Language Modeling for Fast Inference

Namgyu Ho, Sangmin Bae, Taehyeon Kim +6 authors

The Block Transformer architecture enhances inference throughput by applying global-to-local modeling to autoregressive transformers, reducing inference bottlenecks through hierarchical processing and block-level self-attention.

41Block Transformerhierarchical global-to-local modelingHF ↗arXiv ↗
474

sDPO: Don't Use Your Data All at Once

Dahyun Kim, Yungi Kim, Wonho Song +4 authors

A stepwise direct preference optimization approach improves the alignment of large language models with human preferences and enhances their performance.

41large language modelsLLMHF ↗arXiv ↗
476

VideoPrism: A Foundational Visual Encoder for Video Understanding

Long Zhao, Nitesh B. Gundavarapu, Liangzhe Yuan +16 authors

VideoPrism, a pretrained video encoder, achieves top performance across various video understanding tasks by utilizing global-local distillation and token shuffling of semantic video embeddings enhanced with associated text.

41VideoPrismmasked autoencodingHF ↗arXiv ↗
479

SHIC: Shape-Image Correspondences with no Keypoint Supervision

Aleksandar Shtedritski, Christian Rupprecht, Andrea Vedaldi

SHIC leverages foundation computer vision models to learn canonical surface maps without manual supervision, achieving superior results by simulating the annotation process using image-to-image correspondences and enhanced template views.

41DensePosekeypoint detectionHF ↗arXiv ↗
16 / 20

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号