TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

598 篇论文 · 按点赞排序

152

Kvasir-VQA: A Text-Image Pair GI Tract Dataset

Sushant Gautam, Andrea Storås, Cise Midoglu +4 authors

Kvasir-VQA is a dataset with question-and-answer annotations for GI diagnostics, supporting image captioning, VQA, synthetic image generation, object detection, and classification.

71Visual Question Answering (VQA)image captioningHF ↗arXiv ↗
157

BitNet a4.8: 4-bit Activations for 1-bit LLMs

Hongyu Wang, Shuming Ma, Furu Wei

BitNet a4.8 enhances the efficiency of large language models through 4-bit quantization and sparsification, achieving equivalent performance to BitNet b1.58 with reduced inference costs.

701-bit Large Language Models (LLMs)BitNet b1.58HF ↗arXiv ↗
160

Grandmaster-Level Chess Without Search

Anian Ruoss, Grégoire Delétang, Sourabh Medapati +5 authors

A large-scale transformer model trained on a vast dataset of chess games outperforms traditional chess engines and other state-of-the-art models without domain-specific tweaks.

70transformer modelsupervised learningHF ↗arXiv ↗
162

Personalized Visual Instruction Tuning

Renjie Pi, Jianshu Zhang, Tianyang Han +3 authors

A new framework called Personalized Visual Instruction Tuning (PVIT) enhances multimodal large language models to recognize and engage with specific individuals in images, utilizing a curated dataset and benchmarks for evaluation.

70multimodal large language modelsMLLMsHF ↗arXiv ↗
164

Imagine yourself: Tuning-Free Personalized Image Generation

Zecheng He, Bo Sun, Felix Juefei-Xu +14 authors

Imagine yourself is a tuning-free diffusion model for personalized image generation that enhances identity preservation, text alignment, and visual quality through synthetic data generation, parallel attention architecture, and multi-stage fine-tuning.

69diffusion modelspersonalizationHF ↗arXiv ↗
165

Pixtral 12B

Pravesh Agrawal, Szymon Antoniak, Emma Bou Hanna +34 authors

Pixtral-12B, a 12-billion-parameter multimodal language model, excels in both natural language and image understanding, surpassing larger models and introducing an open-source benchmark for evaluation.

69multimodal language modelvision encoderHF ↗arXiv ↗
169

TrustLLM: Trustworthiness in Large Language Models

Lichao Sun, Yue Huang, Haoran Wang +64 authors

This study assesses the trustworthiness of large language models across various dimensions, including truthfulness, safety, fairness, robustness, privacy, and machine ethics, finding a positive correlation with utility and highlighting differences between proprietary and open-source models.

69TrustLLMlarge language modelsHF ↗arXiv ↗
172

Meltemi: The first open Large Language Model for Greek

Leon Voukoutis, Dimitris Roussis, Georgios Paraskevopoulos +6 authors

Developers created Meltemi 7B, a 7 billion parameter open-source large language model for Greek, trained on a 40 billion token corpus, and enhanced with instruction-tuning for chat applications.

68Large Language ModelMistralHF ↗arXiv ↗
174

TÜLU 3: Pushing Frontiers in Open Language Model Post-Training

Nathan Lambert, Jacob Morrison, Valentina Pyatkin +20 authors

T\"ULU 3, an open-source family of post-trained language models, introduces transparent training data, recipes, and advanced techniques to match or surpass proprietary models in performance.

68supervised finetuning (SFT)Direct Preference Optimization (DPO)HF ↗arXiv ↗
175

YuLan-Mini: An Open Data-efficient Language Model

Yiwen Hu, Huatong Song, Jia Deng +8 authors

YuLan-Mini, a 2.42B parameter base model, achieves top-tier performance with efficient pre-training techniques, including a data pipeline with cleaning and scheduling, robust optimization, and annealing with targeted data selection.

67data pipelinedata cleaningHF ↗arXiv ↗
177

LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders

Parishad BehnamGhader, Vaibhav Adlakha, Marius Mosbach +3 authors

LLM2Vec transforms decoder-only LLMs into effective text encoders using bidirectional attention, masked next token prediction, and unsupervised contrastive learning, achieving state-of-the-art performance on text embedding tasks.

67decoder-only language modelsLLM2VecHF ↗arXiv ↗
179

Training-Free Consistent Text-to-Image Generation

Yoad Tewel, Omri Kaduri, Rinon Gal +4 authors

ConsiStory achieves state-of-the-art text-to-image subject consistency without fine-tuning by using shared internal activations, attention blocks, and feature injection.

67text-to-image modelssubject consistencyHF ↗arXiv ↗
6 / 20

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号