TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Sep 23 – Sep 29, 2024
本周最热123

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Multimodal Models

Matt Deitke, Christopher Clark, Sangho Lee +48 authors

Molmo, a new family of open-weight VLMs, excels through a high-quality human-annotated image caption dataset and a diverse fine-tuning dataset, outperforming both proprietary and open models in benchmark and human evaluations.

multimodal modelsopen-weight modelsVLMsimage caption datasetHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

Imagine yourself: Tuning-Free Personalized Image Generation

Zecheng He, Bo Sun, Felix Juefei-Xu +14 authors

Imagine yourself is a tuning-free diffusion model for personalized image generation that enhances identity preservation, text alignment, and visual quality through synthetic data generation, parallel attention architecture, and multi-stage fine-tuning.

69diffusion modelspersonalizationHF ↗arXiv ↗
03

Programming Every Example: Lifting Pre-training Data Quality like Experts at Scale

Fan Zhou, Zengzhi Wang, Qian Liu +2 authors

ProX, a novel framework treating data refinement as a programming task, enhances large language model pre-training by generating fine-grained operations for each example, outperforming traditional human-crafted rule-based methods significantly across various benchmarks and domains.

64Programming Every Example (ProX)data refinementHF ↗arXiv ↗
05

Prithvi WxC: Foundation Model for Weather and Climate

Johannes Schmude, Sujit Roy, Will Trojak +26 authors

Prithvi WxC, a 2.3 billion parameter foundation model using an encoder-decoder architecture with transformer concepts, addresses weather forecasting, downscaling, and extreme events estimation with a mixed masked reconstruction and forecasting objective.

48encoder-decoder-based architecturetransformer modelsHF ↗arXiv ↗
15

EuroLLM: Multilingual Language Models for Europe

Pedro Henrique Martins, Patrick Fernandes, João Alves +12 authors

The EuroLLM project develops multilingual language models for European Union languages, detailing data collection, tokenizer creation, and model performance.

29open-weight LLMsEuroLLM projectHF ↗arXiv ↗
16

Making Text Embedders Few-Shot Learners

Chaofan Li, MingHao Qin, Shitao Xiao +5 authors

A novel model leveraging in-context learning and few-shot examples in large language models improves text embedding generation, achieving state-of-the-art performance on MTEB and AIR-Bench benchmarks.

29Large language modelsdecoder-only architecturesHF ↗arXiv ↗
17

Phantom of Latent for Large Language and Vision Models

Byung-Kwan Lee, Sangyun Chung, Chae Won Kim +2 authors

Phantom, a new efficient LLVM family with smaller model sizes, enhances learning capabilities by temporarily increasing latent hidden dimensions during MHSA and employing Phantom Optimization for improved performance over larger models.

29visual instruction tuninglarge language and vision models (LLVMs)HF ↗arXiv ↗
18

Instruction Following without Instruction Tuning

John Hewitt, Nelson F. Liu, Percy Liang +1 authors

Adaptations without explicit instruction tuning, such as training on responses or narrow-domain instruction-response pairs, can still lead to broad instruction-following behavior in language models.

28instruction tuningimplicit instruction tuningHF ↗arXiv ↗
24

Pixel-Space Post-Training of Latent Diffusion Models

Christina Zhang, Simran Motwani, Matthew Yu +6 authors

Adding pixel-space supervision to latent diffusion models improves high-frequency detail preservation during post-training without compromising text alignment quality.

20latent diffusion modelspixel-space supervisionHF ↗arXiv ↗
25

Present and Future Generalization of Synthetic Image Detectors

Pablo Bernabeu-Perez, Enrique Lopez-Cuena, Dario Garcia-Gasulla

Synthetic image detectors must generalize widely and resist alterations in a rapidly evolving field, where the improvement of generators drives advancements in detectors and vice versa.

20image generation modelssynthetic image detectorsHF ↗arXiv ↗
26

Boosting Healthcare LLMs Through Retrieved Context

Jordi Bayarri-Planas, Ashwin Kumar Gururajan, Dario Garcia-Gasulla

The study enhances the factuality and reliability of large language models in healthcare through optimized context retrieval methods, achieving performance comparable to private solutions in benchmarks and proposing OpenMedPrompt for realistic open-ended answers.

19Large Language Modelscontext retrievalHF ↗arXiv ↗
27

Seeing Faces in Things: A Model and Dataset for Pareidolia

Mark Hamilton, Simon Stent, Vasha DuTell +4 authors

A study investigating the detection of face-like structures in non-face images by both humans and face detectors, revealing differences in pareidolia and proposing a statistical model to explain these detections.

18face pareidoliahuman face detectorHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号