TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jan 1 – Jan 7, 2024
本周最热192

DocLLM: A layout-aware generative language model for multimodal document understanding

Dongsheng Wang, Natraj Raman, Mathieu Sibue +6 authors

DocLLM extends large language models for visual document understanding by focusing on bounding box information and disentangled matrices for multimodal attention, achieving state-of-the-art performance on document intelligence tasks.

multimodal LLMsimage encodersbounding box informationattention mechanismHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

TinyLlama: An Open-Source Small Language Model

Peiyuan Zhang, Guangtao Zeng, Tianduo Wang +1 authors

TinyLlama, a compact 1.1B language model, leverages FlashAttention to achieve high performance in downstream tasks with enhanced computational efficiency.

95FlashAttentionLlama 2HF ↗arXiv ↗
05

Understanding LLMs: A Comprehensive Overview from Training to Inference

Yiheng Liu, Hao He, Tianle Han +18 authors

The paper reviews techniques for cost-efficient training and deployment of large language models, covering aspects like data preprocessing, parallel training, model fine-tuning, and inference optimizations including model compression and memory scheduling.

66Large Language Modelspre-training tasksHF ↗arXiv ↗
06

LLaMA Pro: Progressive LLaMA with Block Expansion

Chengyue Wu, Yukang Gan, Yixiao Ge +5 authors

A new post-pretraining method using expanded Transformer blocks for Large Language Models improves knowledge without catastrophic forgetting, yielding LLaMA Pro-8.3B that excels in general tasks, programming, and mathematics.

54Large Language ModelsLLMsHF ↗arXiv ↗
09

LARP: Language-Agent Role Play for Open-World Games

Ming Yan, Ruihao Li, Hao Zhang +3 authors

LARP is a language agent framework for open-world games, incorporating memory and decision-making capabilities to enhance interaction and coherence in a diverse range of applications.

34cognitive architecturememory processingHF ↗arXiv ↗
11

Instruct-Imagen: Image Generation with Multi-modal Instruction

Hexiang Hu, Kelvin C. K. Chan, Yu-Chuan Su +9 authors

Instruct-Imagen achieves robust and generalized image generation through fine-tuning a pre-trained text-to-image diffusion model with multi-modal instructions and retrieval-augmented training.

31multi-modal instructiontext-to-image diffusion modelHF ↗arXiv ↗
12

aMUSEd: An Open MUSE Reproduction

Suraj Patil, William Berman, Robin Rombach +1 authors

aMUSEd, a masked image model with 10% of MUSE's parameters, generates images quickly and efficiently, requiring fewer inference steps and offering easier fine-tuning compared to latent diffusion.

31masked image modellatent diffusionHF ↗arXiv ↗
16

GPT-4V(ision) is a Generalist Web Agent, if Grounded

Boyuan Zheng, Boyu Gou, Jihyung Kil +2 authors

LMMs like GPT-4V demonstrate potential as generalist web agents by following natural language instructions to complete tasks on live websites, outperforming text-only LLMs and smaller models with a combination of HTML text and visuals for grounding, though challenges remain.

23GPT-4Vmultimodal modelsHF ↗arXiv ↗
17

Astraios: Parameter-Efficient Instruction Tuning Code Large Language Models

Terry Yue Zhuo, Armel Zebaze, Nitchakarn Suppattarachai +4 authors

Astraios, a suite of instruction-tuned models, investigates the cost-performance trade-offs of full-parameter fine-tuning and parameter-efficient fine-tuning across different model sizes and tasks, finding that LoRA offers the best trade-off and FFT generally performs best in downstream tasks.

23full-parameter fine-tuningparameter-efficient fine-tuningHF ↗arXiv ↗
24

COSMO: COntrastive Streamlined MultimOdal Model with Interleaved Pre-Training

Alex Jinpeng Wang, Linjie Li, Kevin Qinghong Lin +5 authors

A contrastive loss and unified multimodal framework improve vision-language pre-training models, particularly in tasks involving long textual contexts and extended video data, outperforming existing models and utilizing a novel interleaved video-text dataset.

17autoregressive vision-language modelscontrastive lossHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号