TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Mar 4 – Mar 10, 2024
本周最热192

GaLore: Memory-Efficient LLM Training by Gradient Low-Rank Projection

Jiawei Zhao, Zhenyu Zhang, Beidi Chen +3 authors

Gradient Low-Rank Projection (GaLore) improves memory efficiency for training large language models without sacrificing performance, enabling pre-training of 7B models on consumer GPUs.

low-rank adaptation (LoRA)Gradient Low-Rank Projection (GaLore)memory efficiencypre-trainingHF ↗arXiv ↗

48 篇论文 · 按点赞排序

03

SaulLM-7B: A pioneering Large Language Model for Law

Pierre Colombo, Telmo Pessoa Pires, Malik Boudiaf +8 authors

SaulLM-7B, a large language model with 7 billion parameters, excels in legal text comprehension and generation using instructional fine-tuning on a legal corpus.

93large language modellegal domainHF ↗arXiv ↗
06

Yi: Open Foundation Models by 01.AI

01. AI, Alex Young, Bei Chen +28 authors

The Yi model family, based on transformer architecture, showcases strong performance across benchmarks and modalities through optimized data and scalable infrastructure.

65language modelsmultimodal modelsHF ↗arXiv ↗
24

Enhancing Vision-Language Pre-training with Rich Supervisions

Yuan Gao, Kunyu Shi, Pengkai Zhu +7 authors

Strongly Supervised pre-training with ScreenShots (S4) enhances Vision-Language Models by using web screenshots with tree-structured HTML elements and spatial localization, resulting in significant improvements in diverse downstream tasks.

17Strongly Supervised pre-trainingScreenShotsHF ↗arXiv ↗
26

Wukong: Towards a Scaling Law for Large-Scale Recommendation

Buyun Zhang, Liang Luo, Yuxin Chen +12 authors

Wukong, a network architecture using stacked factorization machines with a synergistic upscaling strategy, establishes a scaling law in recommendation systems, outperforming state-of-the-art models and maintaining quality across significant increases in complexity.

16scaling lawsfactorization machinesHF ↗arXiv ↗
30

Pix2Gif: Motion-Guided Diffusion for GIF Generation

Hitesh Kandala, Jianfeng Gao, Jianwei Yang

Pix2Gif, a motion-guided diffusion model, generates image-to-GIF videos using text and motion prompts, ensuring coherence and consistency through a perceptual loss and motion-guided warping module.

14motion-guided diffusion modelimage translationHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号