TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Nov 6 – Nov 12, 2023
本周最热85

LCM-LoRA: A Universal Stable-Diffusion Acceleration Module

Simian Luo, Yiqin Tan, Suraj Patil +6 authors

LCMs enhance text-to-image generation by distilling from LDMs using LoRA for reduced memory and superior quality, and introduce LCM-LoRA as a plug-in accelerator for various tasks.

Latent Consistency ModelsLCMslatent diffusion modelsLDMsHF ↗arXiv ↗

42 篇论文 · 按点赞排序

02

LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Shilong Liu, Hao Cheng, Haotian Liu +10 authors

LLaVA-Plus, a general-purpose multimodal assistant, enhances large multimodal models by integrating pre-trained vision and vision-language models, performing tool-assisted tasks and improving interaction through direct image grounding.

50multimodal assistantpre-trained vision and vision-language modelsHF ↗arXiv ↗
03

LRM: Large Reconstruction Model for Single Image to 3D

Yicong Hong, Kai Zhang, Jiuxiang Gu +7 authors

A Large Reconstruction Model using a transformer-based architecture predicts 3D neural radiance fields from single images using massive multi-view training data.

50Large Reconstruction Modeltransformer-based architectureHF ↗arXiv ↗
05

GLaMM: Pixel Grounding Large Multimodal Model

Hanoona Rasheed, Muhammad Maaz, Sahal Shaji +7 authors

GLaMM is a multimodal model that generates visually grounded language responses with object segmentation masks from both text and optional visual prompts.

36Large Multimodal ModelsLarge Language ModelsHF ↗arXiv ↗
06

OtterHD: A High-Resolution Multi-modality Model

Bo Li, Peiyuan Zhang, Jingkang Yang +3 authors

OtterHD-8B, an advanced multimodal model from Fuyu-8B, excels in processing high-resolution inputs and discerning detailed spatial relationships through MagnifierBench, an evaluation framework highlighting the importance of vision encoder flexibility.

34multimodal modelhigh-resolution visual inputsHF ↗arXiv ↗
09

S-LoRA: Serving Thousands of Concurrent LoRA Adapters

Ying Sheng, Shiyi Cao, Dacheng Li +9 authors

S-LoRA is a system that allows for efficient and scalable serving of numerous LoRA adapters using a unified memory pool, tensor parallelism, and custom CUDA kernels.

29Low-Rank Adaptation (LoRA)parameter-efficient fine-tuningHF ↗arXiv ↗
10

CogVLM: Visual Expert for Pretrained Language Models

Weihan Wang, Qingsong Lv, Wenmeng Yu +13 authors

CogVLM, a visual language foundation model, uses a trainable visual expert module to deeply integrate vision and language without compromising NLP performance.

27visual language foundation modelshallow alignment methodHF ↗arXiv ↗
16

Ziya2: Data-centric Learning is All LLMs Need

Ruyi Gan, Ziwei Wu, Renliang Sun +8 authors

Ziya2, a 13-billion-parameter language model built on LLaMA2 and further pre-trained on 700 billion tokens, achieves superior performance across multiple benchmarks using data-centric optimization techniques.

18large language modelsLLMsHF ↗arXiv ↗
17

VR-NeRF: High-Fidelity Virtualized Walkable Spaces

Linning Xu, Vasu Agrawal, William Laney +10 authors

The system captures, reconstructs, and renders high-fidelity walkable spaces using neural radiance fields, achieving high-quality real-time VR rendering with a custom dataset and multi-camera rig.

16neural radiance fieldsmulti-camera rigHF ↗arXiv ↗
18

FLAP: Fast Language-Audio Pre-training

Ching-Feng Yeh, Po-Yao Huang, Vasu Sharma +2 authors

FLAP is a self-supervised method using masking, contrastive learning, and reconstruction to align audio and language representations, achieving state-of-the-art performance in audio-text retrieval tasks.

16self-supervised approachFast Language-Audio Pre-training (FLAP)HF ↗arXiv ↗
21

Holistic Evaluation of Text-To-Image Models

Tony Lee, Michihiro Yasunaga, Chenlin Meng +15 authors

A new benchmark evaluates text-to-image models across 12 aspects, revealing no single model excels comprehensively in all areas.

13text-to-image modelsHolistic Evaluation of Text-to-Image ModelsHF ↗arXiv ↗
22

NExT-Chat: An LMM for Chat, Detection and Segmentation

Ao Zhang, Liming Zhao, Chen-Wei Xie +3 authors

The pixel2emb method enables large multimodal models to effectively handle object location tasks by outputting location embeddings, leading to superior performance in multimodal conversations and various visual processing tasks.

13large language models (LLMs)large multimodal models (LMMs)HF ↗arXiv ↗
24

Can LLMs Follow Simple Rules?

Norman Mu, Sarah Chen, Zifan Wang +5 authors

RuLES is a framework for evaluating rule-following behaviors in LLMs using programmatically defined scenarios and adversarial inputs, revealing vulnerabilities across different models and attack strategies.

12Large Language ModelsLLMsHF ↗arXiv ↗
30

MFTCoder: Boosting Code LLMs with Multitask Fine-Tuning

Bingchang Liu, Chaoyu Chen, Cong Liao +9 authors

MFTcoder, a multi-task fine-tuning framework, improves the coding capabilities of LLMs through parallel fine-tuning on multiple tasks, achieving better performance than single-task fine-tuning and GPT-4 in zero-shot evaluation.

10multi-task fine-tuningMFTcoderHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号