TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

598 篇论文 · 按点赞排序

367

RAPTOR: Recursive Abstractive Processing for Tree-Organized Retrieval

Parth Sarthi, Salman Abdullah, Aditi Tuli +3 authors

The RAPTOR model enhances retrieval-augmented language models by recursively embedding, clustering, and summarizing text chunks, leading to better performance on question-answering tasks involving complex reasoning.

48retrieval-augmented language modelsRecursive embeddingHF ↗arXiv ↗
371

Evaluating and Aligning CodeLLMs on Human Preference

Jian Yang, Jiaxi Yang, Ke Jin +7 authors

A human-curated benchmark (CodeArena) and a large synthetic instruction corpus (SynCode-Instruct) are introduced to evaluate code LLMs based on human preference alignment, revealing performance differences between open-source and proprietary models.

48code large language modelscode generationHF ↗arXiv ↗
372

Lyra: An Efficient and Speech-Centric Framework for Omni-Cognition

Zhisheng Zhong, Chengyao Wang, Yuqi Liu +12 authors

As Multi-modal Large Language Models (MLLMs) evolve, expanding beyond single-domain capabilities is essential to meet the demands for more versatile and efficient AI. However, previous omni-models have insufficiently explored speech, neglecting its integration with multi-modality. We introduce Lyra, an efficient MLLM that enhances multimodal abilities, including advanced long-speech comprehension, sound understanding, cross-modality efficiency, and seamless speech interaction. To achieve efficiency and speech-centric capabilities, Lyra employs three strategies: (1) leveraging existing open-source large models and a proposed multi-modality LoRA to reduce training costs and data requirements; (2) using a latent multi-modality regularizer and extractor to strengthen the relationship between speech and other modalities, thereby enhancing model performance; and (3) constructing a high-quality, extensive dataset that includes 1.5M multi-modal (language, vision, audio) data samples and 12K long speech samples, enabling Lyra to handle complex long speech inputs and achieve more robust omni-cognition. Compared to other omni-methods, Lyra achieves state-of-the-art performance on various vision-language, vision-speech, and speech-language benchmarks, while also using fewer computational resources and less training data.

48Multi-modal Large Language Models (MLLMs)omni-modelsHF ↗arXiv ↗
374

Prithvi WxC: Foundation Model for Weather and Climate

Johannes Schmude, Sujit Roy, Will Trojak +26 authors

Prithvi WxC, a 2.3 billion parameter foundation model using an encoder-decoder architecture with transformer concepts, addresses weather forecasting, downscaling, and extreme events estimation with a mixed masked reconstruction and forecasting objective.

48encoder-decoder-based architecturetransformer modelsHF ↗arXiv ↗
380

FiT: Flexible Vision Transformer for Diffusion Model

Zeyu Lu, Zidong Wang, Di Huang +4 authors

The Flexible Vision Transformer adapts to varied image resolutions and aspect ratios through dynamic tokenization and extrapolation techniques, outperforming traditional methods.

48diffusion modelsDiffusion TransformersHF ↗arXiv ↗
381

Nemotron-4 15B Technical Report

Jupinder Parmar, Shrimai Prabhumoye, Joseph Jennings +24 authors

Nemotron-4 15B, a large multilingual language model, excels in English, multilingual, and coding tasks, demonstrating superior performance in multilingual capabilities compared to larger and specialized models.

48large multilingual language modeldownstream evaluation areasHF ↗arXiv ↗
382

Phased Consistency Model

Fu-Yun Wang, Zhaoyang Huang, Alexander William Bergman +9 authors

The Phased Consistency Model (PCM) addresses limitations in Latent Consistency Models (LCM) and outperforms them in high-resolution, text-conditioned image and few-step text-to-video generation.

48consistency modeldiffusion modelsHF ↗arXiv ↗
390

SemiEvol: Semi-supervised Fine-tuning for LLM Adaptation

Junyu Luo, Xiao Luo, Xiusi Chen +3 authors

A semi-supervised fine-tuning framework named SemiEvol enhances LLM adaptation using both labeled and unlabeled data, showing improved performance through bi-level knowledge propagation and collaborative learning.

47supervised fine-tuninglarge language modelsHF ↗arXiv ↗
13 / 20

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号