TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

50 篇论文 · 按点赞排序

31

Enhancing Human-Like Responses in Large Language Models

Ethem Yağız Çalık, Talha Rüzgar Akkuş

Advancements in enhancing natural language understanding, conversational coherence, and emotional intelligence in large language models improve user interactions and expand AI applications, while future research will address ethical implications and biases.

63large language modelsfine-tuningHF ↗arXiv ↗
34

Baichuan-Omni-1.5 Technical Report

Yadong Li, Jun Liu, Tao Zhang +90 authors

Baichuan-Omni-1.5 is an omni-modal model with end-to-end audio generation, featuring a comprehensive data pipeline, audio-tokenizer, and multi-stage training strategy for superior performance across multimodal tasks.

61omni-modal modelaudio-tokenizerHF ↗arXiv ↗
35

Towards Best Practices for Open Datasets for LLM Training

Stefan Baack, Stella Biderman, Kasia Odrozek +36 authors

Many AI companies are training their large language models (LLMs) on data without the permission of the copyright owners. The permissibility of doing so varies by jurisdiction: in countries like the EU and Japan, this is allowed under certain restrictions, while in the United States, the legal landscape is more ambiguous. Regardless of the legal status, concerns from creative producers have led to several high-profile copyright lawsuits, and the threat of litigation is commonly cited as a reason for the recent trend towards minimizing the information shared about training datasets by both corporate and public interest actors. This trend in limiting data information causes harm by hindering transparency, accountability, and innovation in the broader ecosystem by denying researchers, auditors, and impacted individuals access to the information needed to understand AI models. While this could be mitigated by training language models on open access and public domain data, at the time of writing, there are no such models (trained at a meaningful scale) due to the substantial technical and sociological challenges in assembling the necessary corpus. These challenges include incomplete and unreliable metadata, the cost and complexity of digitizing physical records, and the diverse set of legal and technical skills required to ensure relevance and responsibility in a quickly changing landscape. Building towards a future where AI systems can be trained on openly licensed data that is responsibly curated and governed requires collaboration across legal, technical, and policy domains, along with investments in metadata standards, digitization, and fostering a culture of openness.

61HF ↗arXiv ↗
38

Chain-of-Retrieval Augmented Generation

Liang Wang, Haonan Chen, Nan Yang +3 authors

CoRAG, a multi-step retrieval and reasoning approach, enhances RAG models by dynamically refining queries and using rejection sampling to improve performance, especially in multi-hop question answering.

58RAG modelsCoRAGHF ↗arXiv ↗
42

Transformer^2: Self-adaptive LLMs

Qi Sun, Edoardo Cetin, Yujin Tang

A self-adaptive framework for large language models uses reinforcement learning to dynamically adjust task-specific components during inference, enhancing adaptability and performance with efficiency.

55self-adaptive large language models (LLMs)fine-tuningHF ↗arXiv ↗
45

MinMo: A Multimodal Large Language Model for Seamless Voice Interaction

Qian Chen, Yafeng Chen, Yanni Chen +33 authors

MinMo, a multimodal large language model, integrates speech and text processing to achieve state-of-the-art performance in voice comprehension and generation, while enabling full-duplex conversation and instruction-following capabilities.

54multimodal large language modelnative modelsHF ↗arXiv ↗
50

LTX-Video: Realtime Video Latent Diffusion

Yoav HaCohen, Nisan Chiprut, Benny Brazowski +13 authors

LTX-Video, a transformer-based latent diffusion model, integrates Video-VAE and denoising transformer for efficient high-resolution video generation with temporal consistency.

51transformer-based latent diffusion modelVideo-VAEHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号