TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jun 10 – Jun 16, 2024

50 篇论文 · 按点赞排序

31

Vript: A Video Is Worth Thousands of Words

Dongjie Yang, Suyuan Huang, Chengqiang Lu +5 authors

Vript introduces a high-quality video-text dataset with detailed captions including camera operations, enhancing video captioning and generation, and proposes Vriptor, a top-performing model with Vript-Hard, a new benchmark for video understanding challenges.

26multimodal learningvideo understandingHF ↗arXiv ↗
34

Towards a Personal Health Large Language Model

Justin Cosentino, Anastasiya Belyaeva, Xin Liu +31 authors

A Personal Health Large Language Model (PH-LLM) fine-tuned from Gemini demonstrates substantial capabilities in processing and interpreting personal health data, offering personalized insights and outperforming experts in certain domain-specific tasks.

21Large Language Model (LLM)Personal Health Large Language Model (PH-LLM)HF ↗arXiv ↗
35

Large Language Model Confidence Estimation via Black-Box Access

Tejaswini Pedapati, Amit Dhurandhar, Soumya Ghosh +2 authors

A framework using engineered features and logistic regression effectively estimates the confidence of large language models with black-box access, outperforming existing methods on multiple datasets.

21large language modelsblack-box accessHF ↗arXiv ↗
36

GenAI Arena: An Open Evaluation Platform for Generative Models

Dongfu Jiang, Max Ku, Tianle Li +4 authors

GenAI-Arena, an open platform, leverages user feedback to evaluate generative models across text-to-image, text-to-video, and image editing tasks, demonstrating existing multimodal models' shortcomings in assessing generated content quality.

21generative AIimage generationHF ↗arXiv ↗
37

Interpreting the Weight Space of Customized Diffusion Models

Amil Dravid, Yossi Gandelsman, Kuan-Chieh Wang +4 authors

A large dataset of fine-tuned diffusion models for different visual identities reveals a manifold representing identity in the weight space, enabling sampling, semantic editing, and inversion tasks.

20diffusion modelsweights2weightsHF ↗arXiv ↗
43

Tx-LLM: A Large Language Model for Therapeutics

Juan Manuel Zambrano Chaves, Eric Wang, Tao Tu +7 authors

Tx-LLM, a specialized large language model fine-tuned from PaLM-2, encodes diverse therapeutic knowledge and achieves superior performance across various drug discovery tasks by interleaving chemical data with free-text.

18large language modelLLMHF ↗arXiv ↗
49

SEE-2-SOUND: Zero-Shot Spatial Environment-to-Spatial Sound

Rishit Dagli, Shivesh Prakash, Robert Wu +1 authors

SEE-2-SOUND generates spatial audio for visual content by decomposing the task into identifying visual regions, locating them in 3D space, generating mono-audio for each, and integrating it into spatial audio, addressing a gap in current audio generation models.

15neural generative modelshigh-resolution contentHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号