TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

598 篇论文 · 按点赞排序

241

AniDoc: Animation Creation Made Easier

Yihao Meng, Hao Ouyang, Hanlin Wang +6 authors

AniDoc uses video diffusion models to automate colorization and in-betweening in 2D animation, improving efficiency by leveraging correspondence matching.

58video diffusion modelscorrespondence matchingHF ↗arXiv ↗
248

More Agents Is All You Need

Junyou Li, Qin Zhang, Yangbin Yu +2 authors

A sampling-and-voting method enhances large language models' performance by increasing the number of agents, with effectiveness tied to task difficulty.

57large language modelsLLMSHF ↗arXiv ↗
251

Aya Dataset: An Open-Access Collection for Multilingual Instruction Tuning

Shivalika Singh, Freddie Vargus, Daniel Dsouza +30 authors

The initiative builds a human-curated instruction-following dataset spanning 65 languages and creates the largest multilingual collection of instruction-following instances through templating and translating existing datasets across 114 languages, contributing datasets and platforms for participatory research.

57Instruction fine-tuningIFTHF ↗arXiv ↗
264

ViTAR: Vision Transformer with Any Resolution

Qihang Fan, Quanzeng You, Xiaotian Han +5 authors

ViTAR enhances Vision Transformers' scalability across resolutions through dynamic token integration and fuzzy positional encoding, improving accuracy and reducing computational costs.

56Vision Transformersdynamic resolution adjustmentHF ↗arXiv ↗
265

CoD, Towards an Interpretable Medical Agent using Chain of Diagnosis

Junying Chen, Chi Gui, Anningzhe Gao +4 authors

Chain-of-Diagnosis (CoD) enhances interpretability in LLM-based medical diagnostics by providing a transparent reasoning pathway and developing DiagnosisGPT, which diagnoses a wide range of diseases with high accuracy and controllable rigor.

55large language models (LLMs)Chain-of-Diagnosis (CoD)HF ↗arXiv ↗
267

Transformers Can Do Arithmetic with the Right Embeddings

Sean McLeish, Arpit Bansal, Alex Stein +8 authors

Transformers achieve state-of-the-art performance on large arithmetic tasks and other reasoning tasks by addressing positional tracking with embeddings and integrating architectural modifications.

55transformerspositional trackingHF ↗arXiv ↗
268

Meteor: Mamba-based Traversal of Rationale for Large Language and Vision Models

Byung-Kwan Lee, Chae Won Kim, Beomchan Park +1 authors

Meteor, an efficient large language and vision model, enhances understanding and answering capabilities by embedding multifaceted rationales using the Mamba architecture, leading to improved vision language performance without increasing model size or using additional vision encoders.

55large language and vision modelsvisual instruction tuningHF ↗arXiv ↗
269

Needle In A Multimodal Haystack

Weiyun Wang, Shuibo Zhang, Yiming Ren +13 authors

NN-NIAH is a benchmark that evaluates MLLMs' comprehension of long multimodal documents across retrieval, counting, and reasoning tasks.

55multimodal large language modelsMLLMsHF ↗arXiv ↗
9 / 20

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号