TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jul 31 – Aug 6, 2023
本周最热102

ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs

Yujia Qin, Shihao Liang, Yining Ye +15 authors

ToolLLM is a framework that enhances open-source LLMs with tool-use capabilities through a comprehensive dataset, efficient decision tree, automatic evaluation, and neural API retrieval.

instruction tuningtool-use frameworkToolBenchRESTful APIsHF ↗arXiv ↗

42 篇论文 · 按点赞排序

03

LP-MusicCaps: LLM-Based Pseudo Music Captioning

SeungHeon Doh, Keunwoo Choi, Jongpil Lee +1 authors

A large-scale pseudo music caption dataset generated using LLMs improves music captioning performance over supervised models in zero-shot and transfer-learning settings.

39large language modelsLLMsHF ↗arXiv ↗
07

Learning to Model the World with Language

Jessy Lin, Yuqing Du, Olivia Watkins +4 authors

Dynalang, a multimodal agent, learns to predict future text and image representations using language hints to improve task performance and enrich its understanding.

36multimodal world modelself-supervised learningHF ↗arXiv ↗
11

Med-Flamingo: a Multimodal Medical Few-shot Learner

Michael Moor, Qian Huang, Shirley Wu +6 authors

Med-Flamingo, an adaptation of OpenFlamingo-9B for the medical domain, demonstrates few-shot capabilities in generative visual question answering with significant performance improvements as evaluated by clinicians.

25medical generative vision-language modelsfew-shot learnerHF ↗arXiv ↗
15

From Sparse to Soft Mixtures of Experts

Joan Puigcerver, Carlos Riquelme, Basil Mustafa +1 authors

Soft MoE, a differentiable sparse Transformer, stabilizes training, reduces inference cost, and outperforms traditional Transformers and MoE variants in visual recognition.

22sparse mixture of expert architecturesMoEsHF ↗arXiv ↗
18

Guiding Image Captioning Models Toward More Specific Captions

Simon Kornblith, Lala Li, Zirui Wang +1 authors

Classifier-free guidance improves specific caption generation by fine-tuning for both conditional and unconditional distributions, enhancing reference-free metrics and web data quality.

18autoregressive captioning modelclassifier-free guidanceHF ↗arXiv ↗
19

Multimodal Neurons in Pretrained Text-Only Transformers

Sarah Schwettmann, Neil Chowdhury, Antonio Torralba

Frozen text transformers can incorporate visual information through self-supervised learning, with specific neurons facilitating translation between modalities and influencing image captioning.

17text transformerself-supervised visual encoderHF ↗arXiv ↗
22

Unified Model for Image, Video, Audio and Language Tasks

Mustafa Shukor, Corentin Dancette, Alexandre Rame +1 authors

UnIVAL, a unified model with 0.25B parameters, efficiently supports text, images, video, and audio, demonstrating competitive performance and benefiting out-of-distribution generalization through multimodal weight interpolation.

16Large Language Modelsgeneralist agentsHF ↗arXiv ↗
27

ELIXR: Towards a general purpose X-ray artificial intelligence system through alignment of large language models and radiology vision encoders

Shawn Xu, Lin Yang, Christopher Kelly +25 authors

ELIXR, a lightweight adapter architecture combining a language-aligned image encoder with PaLM 2, achieves state-of-the-art performance in zero-shot and data-efficient chest X-ray classification and semantic search, requiring significantly less data compared to existing methods.

13language-aligned image encoderLLMHF ↗arXiv ↗
30

UniVTG: Towards Unified Video-Language Temporal Grounding

Kevin Qinghong Lin, Pengchuan Zhang, Joya Chen +5 authors

UniVTG unifies various video temporal grounding tasks and labels through a unified framework, enabling effective and flexible model development and delivering strong performance across multiple datasets.

12Video Temporal Grounding (VTG)moment retrievalHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号