TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Feb 3 – Feb 9, 2025

50 篇论文 · 按点赞排序

31

Reformulation for Pretraining Data Augmentation

Xintong Hao, Ruijie Zhu, Ge Zhang +2 authors

The Massive Genre-Audience reformulation method augments training data to reduce repetition-related degradation and supports more efficient scaling of large language models.

23Massive Genre-Audience reformulationsynthetic data augmentationHF ↗arXiv ↗
41

On Teacher Hacking in Language Model Distillation

Daniil Tiapkin, Daniele Calandriello, Johan Ferret +4 authors

Experiments reveal that knowledge distillation can lead to teacher hacking, an issue akin to reward hacking, but it can be mitigated by using online data generation techniques that enhance data diversity.

19knowledge distillationreinforcement learning from human feedbackHF ↗arXiv ↗
42

AIN: The Arabic INclusive Large Multimodal Model

Ahmed Heakl, Sara Ghaboura, Omkar Thawkar +4 authors

AIN, a bilingual English-Arabic multimodal model, achieves state-of-the-art performance in Arabic and strong English-language visual capabilities across diverse domains.

19large language modelslarge multimodal modelsHF ↗arXiv ↗
47

PixelWorld: Towards Perceiving Everything as Pixels

Zhiheng Lyu, Xueguang Ma, Wenhu Chen

A unified perception framework treating all modalities as pixel inputs (PEAP) outperforms token-based input in some tasks but degrades reasoning and coding capabilities, requiring enhancements in foundational models’ perceptual abilities.

15PixelWorldmultimodal datasetsHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号