TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

November 2024

50 篇论文 · 按点赞排序

32

Cut Your Losses in Large-Vocabulary Language Models

Erik Wijmans, Brody Huval, Alexander Hertzberg +2 authors

Cut Cross-Entropy reduces the memory footprint of large language models during training by efficiently computing cross-entropy loss without materializing logits for all tokens.

49Cross-EntropyCut Cross-Entropy (CCE)HF ↗arXiv ↗
39

Style-Friendly SNR Sampler for Style-Driven Generation

Jooyoung Choi, Chaehun Shin, Yeongtak Oh +2 authors

The Style-friendly SNR sampler modifies the noise level distribution during fine-tuning to improve style alignment in diffusion models, enabling better capture of unique artistic styles.

40diffusion modelssignal-to-noise ratio (SNR)HF ↗arXiv ↗
43

Stronger Models are NOT Stronger Teachers for Instruction Tuning

Zhangchen Xu, Fengqing Jiang, Luyao Niu +2 authors

The Larger Models' Paradox reveals that larger models are not always better teachers for fine-tuning smaller models, and a new metric, Compatibility-Adjusted Reward (CAR), is introduced to measure and improve the effectiveness of response generators.

39instruction tuninglarge language models (LLMs)HF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号