TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Mar 18 – Mar 24, 2024
本周最热187

LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Yaowei Zheng, Richong Zhang, Junhao Zhang +2 authors

LlamaFactory is a unified framework enabling efficient fine-tuning of large language models across various tasks using a web-based user interface.

efficient fine-tuninglarge language modelsLLaMALlamaFactoryHF ↗arXiv ↗

50 篇论文 · 按点赞排序

03

RAFT: Adapting Language Model to Domain Specific RAG

Tianjun Zhang, Shishir G. Patil, Naman Jain +4 authors

Retrieval Augmented FineTuning (RAFT) enhances pre-trained large language models' in-domain performance by filtering out irrelevant documents and citing relevant passages.

72Retrieval Augmented FineTuningRAFTHF ↗arXiv ↗
06

Evolutionary Optimization of Model Merging Recipes

Takuya Akiba, Makoto Shing, Yujin Tang +2 authors

An evolutionary algorithm automates model merging, discovering effective combinations of open-source models without additional training data or compute, achieving state-of-the-art performance on benchmarks and culturally-aware tasks.

58evolutionary algorithmsfoundation modelsHF ↗arXiv ↗
13

mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Anwen Hu, Haiyang Xu, Jiabo Ye +8 authors

A Unified Structure Learning framework is introduced to enhance Visual Document Understanding by integrating structure-aware parsing and multi-grained text localization tasks using a vision-to-text module H-Reducer, achieving state-of-the-art results on multiple benchmarks.

32Multimodal Large Language ModelsVisual Document UnderstandingHF ↗arXiv ↗
17

When Do We Not Need Larger Vision Models?

Baifeng Shi, Ziyang Wu, Maolin Mao +2 authors

Running smaller vision models at multiple image scales can achieve state-of-the-art performance on various tasks, often surpassing larger models.

26Scaling on Scales (S$^2$)pre-trained vision modelHF ↗arXiv ↗
22

TnT-LLM: Text Mining at Scale with Large Language Models

Mengting Wan, Tara Safavi, Sujay Kumar Jauhar +11 authors

TnT-LLM is a two-phase framework using Large Language Models to automate label generation and classification with minimal human effort, demonstrating superior accuracy and efficiency in text mining tasks.

21Large Language Modelsprompt-based interfaceHF ↗arXiv ↗
29

DepthFM: Fast Monocular Depth Estimation with Flow Matching

Ming Gui, Johannes S. Fischer, Ulrich Prestel +6 authors

A flow matching approach using pre-trained image diffusion models achieves state-of-the-art monocular depth estimation with high quality and low computational cost, leveraging synthetic data and surface normals loss.

18monocular depth estimationdiscriminative approachesHF ↗arXiv ↗
30

ZigMa: Zigzag Mamba Diffusion Model

Vincent Tao Hu, Stefan Andreas Baumann, Ming Gui +4 authors

Zigzag Mamba, an improvement over Mamba, enhances visual data generation with better spatial continuity, speed, and memory usage, and is scalable with Stochastic Interpolant on large-resolution datasets.

18diffusion modelscalabilityHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号