TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

598 篇论文 · 按点赞排序

334

LTX-Video: Realtime Video Latent Diffusion

Yoav HaCohen, Nisan Chiprut, Benny Brazowski +13 authors

LTX-Video, a transformer-based latent diffusion model, integrates Video-VAE and denoising transformer for efficient high-resolution video generation with temporal consistency.

51transformer-based latent diffusion modelVideo-VAEHF ↗arXiv ↗
335

DeepSeek-VL: Towards Real-World Vision-Language Understanding

Haoyu Lu, Wen Liu, Bo Zhang +11 authors

DeepSeek-VL is an open-source vision-language model that achieves state-of-the-art performance in real-world applications by combining a hybrid vision encoder with effective pretraining strategies to preserve language model capabilities.

50DeepSeek-VLVision-Language ModelHF ↗arXiv ↗
336

Hymba: A Hybrid-head Architecture for Small Language Models

Xin Dong, Yonggan Fu, Shizhe Diao +10 authors

Hymba, a family of small language models with a hybrid-head architecture combining transformer attention and state space models, achieves state-of-the-art performance with improved efficiency and reduced cache size.

50hybrid-head parallel architecturetransformer attention mechanismsHF ↗arXiv ↗
343

Iterative Reasoning Preference Optimization

Richard Yuanzhe Pang, Weizhe Yuan, Kyunghyun Cho +3 authors

An iterative preference optimization method using a modified DPO loss improves reasoning accuracy on various datasets by optimizing winning and losing reasoning steps in Chain-of-Thought candidates.

50iterative preference optimizationChain-of-Thought (CoT)HF ↗arXiv ↗
347

MoRA: High-Rank Updating for Parameter-Efficient Fine-Tuning

Ting Jiang, Shaohan Huang, Shengyue Luo +8 authors

MoRA, a high-rank updating method using square matrices, enhances the ability of large language models to learn and memorize new knowledge, especially in memory-intensive tasks, compared to LoRA.

50low-rank adaptationparameter-efficient fine-tuningHF ↗arXiv ↗
348

Chronos: Learning the Language of Time Series

Abdul Fatir Ansari, Lorenzo Stella, Caner Turkmen +14 authors

Chronos, a framework using pretrained transformer-based models for time series forecasting, outperforms classical methods and achieves comparable zero-shot performance on unseen datasets using tokenized time series data.

50pretrained probabilistic time series modelsChronosHF ↗arXiv ↗
349

VITA: Towards Open-Source Interactive Omni Multimodal LLM

Chaoyou Fu, Haojia Lin, Zuwei Long +12 authors

VITA, an open-source Multimodal Large Language Model, excels in processing Video, Image, Text, and Audio with seamless interaction, showcasing advancements in multimodal understanding and human-computer interaction.

50Multimodal Large Language ModelMixtralHF ↗arXiv ↗
351

Multimodal Latent Language Modeling with Next-Token Diffusion

Yutao Sun, Hangbo Bao, Wenhui Wang +5 authors

LatentLM integrates continuous and discrete data using causal Transformers, VAEs, and next-token diffusion, achieving superior performance in multimodal tasks like image generation, large language model integration, and text-to-speech synthesis.

50LatentLMcausal TransformersHF ↗arXiv ↗
352

MMEvol: Empowering Multimodal Large Language Models with Evol-Instruct

Run Luo, Haonan Zhang, Longze Chen +13 authors

MMEvol, a multimodal instruction data evolution framework, enhances the capabilities of Multimodal Large Language Models by generating diverse and complex image-text instruction datasets, leading to improved performance across various vision-language tasks.

49Multimodal Large Language ModelsMMEvolHF ↗arXiv ↗
353

TabReD: A Benchmark of Tabular Machine Learning in-the-Wild

Ivan Rubachev, Nikolay Kartashev, Yury Gorishniy +1 authors

TabReD is a new collection of industry-grade tabular datasets that address temporal dynamics and feature engineering, revealing that simpler models outperform more complex deep learning architectures in these settings.

49TabReDMLP-like architecturesHF ↗arXiv ↗
354

MagicVideo-V2: Multi-Stage High-Aesthetic Video Generation

Weimin Wang, Jiawei Liu, Zhijie Lin +9 authors

MagicVideo-V2 generates high-fidelity and smooth videos from text using an integrated pipeline that includes text-to-image, video motion generation, and frame interpolation modules, outperforming existing systems in user evaluations.

49text-to-imagevideo motion generatorHF ↗arXiv ↗
355

Video Diffusion Alignment via Reward Gradients

Mihir Prabhudesai, Russell Mendonca, Zheyang Qin +2 authors

Utilizing pre-trained reward models to adapt video diffusion models with gradient-based feedback enhances efficiency and performance compared to gradient-free methods.

49video diffusion modelspre-trained reward modelsHF ↗arXiv ↗
358

Cut Your Losses in Large-Vocabulary Language Models

Erik Wijmans, Brody Huval, Alexander Hertzberg +2 authors

Cut Cross-Entropy reduces the memory footprint of large language models during training by efficiently computing cross-entropy loss without materializing logits for all tokens.

49Cross-EntropyCut Cross-Entropy (CCE)HF ↗arXiv ↗
12 / 20

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号