TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

本月最热176

Qwen2 Technical Report

An Yang, Baosong Yang, Binyuan Hui +55 authors

The Qwen2 series, comprising 0.5 to 72 billion parameter models, surpasses prior open models across language understanding, generation, multilingualism, coding, math, and reasoning, with exceptional performance in benchmarks like MMLU, GPQA, HumanEval, GSM8K, BBH, MT-Bench, Arena-Hard, and LiveCodeBench.

Mixture-of-Expertslanguage modelsmultimodal modelsMMLUHF ↗arXiv ↗

50 篇论文 · 按点赞排序

07

Vision language models are blind

Pooyan Rahmanzadehgervi, Logan Bolton, Mohammad Reza Taesiri +1 authors

State-of-the-art large language models with vision capabilities perform poorly on simple visual tasks, indicating limitations in their visual understanding.

84large language modelsvision capabilitiesHF ↗arXiv ↗
10

PaliGemma: A versatile 3B VLM for transfer

Lucas Beyer, Andreas Steiner, André Susano Pinto +32 authors

PaliGemma, a versatile Vision-Language Model based on SigLIP-So400m and Gemma-2B, demonstrates strong performance across numerous open-world tasks, including specialized areas like remote sensing and segmentation.

73Vision-Language ModelSigLIP-So400mHF ↗arXiv ↗
11

Meltemi: The first open Large Language Model for Greek

Leon Voukoutis, Dimitris Roussis, Georgios Paraskevopoulos +6 authors

Developers created Meltemi 7B, a 7 billion parameter open-source large language model for Greek, trained on a 40 billion token corpus, and enhanced with instruction-tuning for chat applications.

68Large Language ModelMistralHF ↗arXiv ↗
12

SaulLM-54B & SaulLM-141B: Scaling Up Domain Adaptation for the Legal Domain

Pierre Colombo, Telmo Pires, Malik Boudiaf +7 authors

Two large legal language models, SaulLM-54B and SaulLM-141B, based on the Mixtral architecture, are introduced for domain-specific adaptation in the legal sector using continued pretraining, specialized protocols, and preference alignment with synthetic data.

66Mixtral architecturelarge language modelsHF ↗arXiv ↗
14

Qwen2-Audio Technical Report

Yunfei Chu, Jin Xu, Qian Yang +9 authors

Qwen2-Audio, a large-scale audio-language model, enhances instruction-following and audio analysis through natural language prompts and DPO optimization.

64audio-language modelpre-training processHF ↗arXiv ↗
23

CoD, Towards an Interpretable Medical Agent using Chain of Diagnosis

Junying Chen, Chi Gui, Anningzhe Gao +4 authors

Chain-of-Diagnosis (CoD) enhances interpretability in LLM-based medical diagnostics by providing a transparent reasoning pathway and developing DiagnosisGPT, which diagnoses a wide range of diseases with high accuracy and controllable rigor.

55large language models (LLMs)Chain-of-Diagnosis (CoD)HF ↗arXiv ↗
25

Unveiling Encoder-Free Vision-Language Models

Haiwen Diao, Yufeng Cui, Xiaotong Li +3 authors

EVE is an encoder-free vision-language model that achieves competitive performance on multiple benchmarks using a unified decoder and extra supervision.

55vision-language modelsVLMsHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号