TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Feb 24 – Mar 2, 2025
本周最热175

LLM-Microscope: Uncovering the Hidden Role of Punctuation in Context Memory of Transformers

Anton Razzhigaev, Matvey Mikhalchuk, Temurbek Rahmatullaev +4 authors

Analysis of Large Language Models reveals that even minor tokens are crucial for context, and introduces LLM-Microscope for assessing token-level nonlinearity and contextual memory.

Large Language ModelsLLMscontextual informationtokensHF ↗arXiv ↗

50 篇论文 · 按点赞排序

03

Self-rewarding correction for mathematical reasoning

Wei Xiong, Hanning Zhang, Chenlu Ye +3 authors

Self-rewarding reasoning large language models independently generate and correct their outputs during inference using a two-stage algorithmic framework, enhancing performance without external feedback.

82self-rewarding reasoninglarge language modelsHF ↗arXiv ↗
07

Thus Spake Long-Context Large Language Model

Xiaoran Liu, Ruixiao Li, Mianqiu Huang +10 authors

The survey examines the advancements and challenges in long-context Large Language Models (LLMs), exploring the architecture, infrastructure, training, and evaluation technologies needed to extend their context length and address the inherent trade-offs.

73Large Language Modelslong contextHF ↗arXiv ↗
09

GHOST 2.0: generative high-fidelity one shot transfer of heads

Alexander Groshev, Anastasiia Iashchenko, Pavel Paramonov +2 authors

GHOST 2.0, consisting of an Aligner and Blender module, achieves state-of-the-art results in head swapping by preserving identity information, handling extreme poses, and seamlessly integrating the reenacted head into the target background.

67AlignerBlenderHF ↗arXiv ↗
10

Kanana: Compute-efficient Bilingual Language Models

Kanana LLM Team, Yunju Bak, Hojin Lee +26 authors

Kanana, a series of bilingual language models, achieves superior performance in Korean and competitive performance in English with lower computational costs through efficient pre-training and post-training techniques.

66high quality data filteringstaged pre-trainingHF ↗arXiv ↗
13

Towards an AI co-scientist

Juraj Gottweis, Wei-Hung Weng, Alexander Daryin +31 authors

A multi-agent AI system named AI co-scientist aids in scientific discovery by generating and validating novel hypotheses across biomedical areas, demonstrating potential improvements in drug repurposing, target discovery, and bacterial evolution understanding.

54multi-agent systemGemini 2.0HF ↗arXiv ↗
14

DICEPTION: A Generalist Diffusion Model for Visual Perceptual Tasks

Canyu Zhao, Mingyu Liu, Huanyi Zheng +5 authors

DICEPTION, a text-to-image diffusion model, achieves state-of-the-art performance in multiple perception tasks using minimal data and computational resources by leveraging color encoding and conditional image generation.

51text-to-image diffusion modelsperception tasksHF ↗arXiv ↗
19

NeoBERT: A Next-Generation BERT

Lola Le Breton, Quentin Fournier, Mariam El Mezouar +1 authors

NeoBERT is a next-generation encoder that integrates modern advancements to surpass performance of BERT and RoBERTa on MTEB with a compact parameter footprint.

39NeoBERTstate-of-the-art advancementsHF ↗arXiv ↗
23

Audio-FLAN: A Preliminary Release

Liumeng Xue, Ziya Zhou, Jiahao Pan +19 authors

A new large-scale dataset, Audio-FLAN, supports unified audio-language models by covering diverse tasks across speech, music, and sound for both understanding and generation.

36audio tokenizationlarge language models (LLMs)HF ↗arXiv ↗
28

Plutus: Benchmarking Large Language Models in Low-Resource Greek Finance

Xueqing Peng, Triantafillos Papadopoulos, Efstathia Soufleri +7 authors

The first Greek Financial Evaluation Benchmark (Plutus-ben) and Greek Financial LLM (Plutus-8B) address the challenges of financial NLP in Greek due to linguistic complexity and lack of domain-specific data.

32large language modelsmultilingual financial natural language processingHF ↗arXiv ↗
29

SIFT: Grounding LLM Reasoning in Contexts via Stickers

Zihao Zeng, Xuyao Huang, Boxiu Li +1 authors

A novel post-training approach, Stick to the Facts (SIFT), improves the reasoning accuracy of large language models by generating and refining stickers, leading to enhanced performance on diverse benchmarks.

32Large language modelsStick to the Facts (SIFT)HF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号