TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Dec 23 – Dec 29, 2024
本周最热87

RobustFT: Robust Supervised Fine-tuning for Large Language Models under Noisy Response

Junyu Luo, Xiao Luo, Kaize Ding +3 authors

A noise-robust supervised fine-tuning framework enhances large language models' performance in downstream tasks by detecting and relabeling noisy data.

supervised fine-tuninglarge language modelsnoise-robust frameworkmulti-expert collaborative systemHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

YuLan-Mini: An Open Data-efficient Language Model

Yiwen Hu, Huatong Song, Jia Deng +8 authors

YuLan-Mini, a 2.42B parameter base model, achieves top-tier performance with efficient pre-training techniques, including a data pipeline with cleaning and scheduling, robust optimization, and annealing with targeted data selection.

66data pipelinedata cleaningHF ↗arXiv ↗
03

Parallelized Autoregressive Visual Generation

Yuqing Wang, Shuhuai Ren, Zhijie Lin +6 authors

A parallel generation strategy for autoregressive models improves inference speed without significantly compromising quality in visual generation tasks.

52autoregressive modelsparallelized autoregressiveHF ↗arXiv ↗
05

Token-Budget-Aware LLM Reasoning

Tingxu Han, Chunrong Fang, Shiyu Zhao +3 authors

A token-budget-aware framework dynamically estimates and allocates token budgets for LLM reasoning, reducing costs with minimal performance loss.

45Chain-of-ThoughtCoT reasoningHF ↗arXiv ↗
06

Diving into Self-Evolving Training for Multimodal Reasoning

Wei Liu, Junlong Li, Xiwen Zhang +3 authors

A study on self-evolving training for multimodal reasoning identifies key factors and configurations to optimize reasoning abilities without additional human annotations, introducing MSTaR as an effective training framework.

41Large Multimodal Modelsself-evolving trainingHF ↗arXiv ↗
09

OpenAI o1 System Card

OpenAI, Aaron Jaech, Adam Kalai +262 authors

The o1 model series, trained with reinforcement learning and chain of thought, enhances safety and robustness by reasoning about policies, leading to superior performance on risk benchmarks while highlighting the need for robust alignment and risk management.

38reinforcement learningchain of thoughtHF ↗arXiv ↗
10

Offline Reinforcement Learning for LLM Multi-Step Reasoning

Huaijie Wang, Shibo Hao, Hanze Dong +4 authors

OREO, an offline reinforcement learning method, enhances multi-step reasoning in large language models by optimizing a policy model and value function, improving performance on benchmarks without the need for paired preference data.

38Direct Preference Optimization (DPO)offline reinforcement learning (RL)HF ↗arXiv ↗
13

DepthLab: From Partial to Complete

Zhiheng Liu, Ka Leong Cheng, Qiuyu Wang +7 authors

DepthLab, a depth inpainting model utilizing image diffusion priors, effectively handles missing depth data, preserving scale consistency and excelling in tasks like 3D scene inpainting and LiDAR depth completion.

36depth inpaintingimage diffusion priorsHF ↗arXiv ↗
17

Large Motion Video Autoencoding with Cross-modal Video VAE

Yazhou Xing, Yang Fei, Yingqing He +4 authors

A novel video Variational Autoencoder (VAE) integrates temporal-aware spatial compression and motion compression with text guidance to achieve high-fidelity video encoding and superior reconstruction quality.

24Video VAEtemporal compressionHF ↗arXiv ↗
18

LearnLM: Improving Gemini for Learning

LearnLM Team, Abhinit Modi, Aditya Srikanth Veerubhotla +43 authors

Pedagogical instruction following enhances generative AI systems for educational purposes, leading to a preferred model (LearnLM) across various learning scenarios.

23pedagogical instruction followingLearnLMHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号