TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

403 篇论文 · 按点赞排序

211

Collaborative Score Distillation for Consistent Visual Synthesis

Subin Kim, Kyungmin Lee, June Suk Choi +3 authors

A novel method, Collaborative Score Distillation (CSD), based on Stein Variational Gradient Descent (SVGD), enhances consistency in text-to-image diffusion models across multiple images such as panoramas, videos, and 3D scenes.

31text-to-image diffusion modelsCollaborative Score Distillation (CSD)HF ↗arXiv ↗
214

SLiMe: Segment Like Me

Aliasghar Khani, Saeid Asgari Taghanaki, Aditya Sanghi +2 authors

SLiMe segments images at desired granularity using Stable Diffusion with minimal annotations and outperforms existing one-shot and few-shot segmentation methods.

31Stable Diffusionattention mapsHF ↗arXiv ↗
217

Fine-tuning Language Models for Factuality

Katherine Tian, Eric Mitchell, Huaxiu Yao +2 authors

Fine-tuning language models using automatically generated factuality preference rankings improves their factual accuracy without human labeling.

30large pre-trained language modelsLLMsHF ↗arXiv ↗
222

VeRA: Vector-based Random Matrix Adaptation

Dawid Jan Kopiczko, Tijmen Blankevoort, Yuki Markus Asano

Vector-based Random Matrix Adaptation (VeRA) reduces the number of trainable parameters by 10x compared to LoRA while maintaining performance, and is demonstrated on benchmarks like GLUE and E2E, showing its utility in instruction-following.

30Low-rank adaptationLoRAHF ↗arXiv ↗
223

Secrets of RLHF in Large Language Models Part I: PPO

Rui Zheng, Shihan Dou, Songyang Gao +24 authors

This report examines Reinfocement Learning with Human Feedback (RLHF) and proposes PPO-max to improve the stability of policy model training compared to other SFT models and ChatGPT.

30Reinfocement Learning with Human Feedback (RLHF)reward modelsHF ↗arXiv ↗
225

FinGPT: Large Generative Models for a Small Language

Risto Luukkonen, Ville Komulainen, Jouni Luoma +18 authors

The study addresses the challenges of creating large language models for underrepresented languages like Finnish, developing both monolingual and multilingual models, and evaluating their performance through a newly created benchmark.

30large language modelsLLMsHF ↗arXiv ↗
227

S-LoRA: Serving Thousands of Concurrent LoRA Adapters

Ying Sheng, Shiyi Cao, Dacheng Li +9 authors

S-LoRA is a system that allows for efficient and scalable serving of numerous LoRA adapters using a unified memory pool, tensor parallelism, and custom CUDA kernels.

30Low-Rank Adaptation (LoRA)parameter-efficient fine-tuningHF ↗arXiv ↗
231

AlphaStar Unplugged: Large-Scale Offline Reinforcement Learning

Michaël Mathieu, Sherjil Ozair, Srivatsan Srinivasan +21 authors

A benchmark called AlphaStar Unplugged is established for offline reinforcement learning using a dataset from StarCraft II, featuring unprecedented challenges and achieving high win rates with offline data.

29offline RL algorithmsbehavior cloningHF ↗arXiv ↗
235

Stay on topic with Classifier-Free Guidance

Guillaume Sanchez, Honglu Fan, Alexander Spangher +3 authors

Classifier-Free Guidance enhances performance across various language modeling tasks and improves the faithfulness and coherence of AI assistants, outperforming models with higher parameter counts.

29Classifier-Free GuidancePythiaHF ↗arXiv ↗
236

Tracking Anything with Decoupled Video Segmentation

Ho Kei Cheng, Seoung Wug Oh, Brian Price +2 authors

DEVA, a decoupled approach using image-level segmentation and bi-directional temporal propagation, achieves favorable results in various data-scarce video segmentation tasks with reduced annotation and training costs.

29video segmentationtask-specificHF ↗arXiv ↗
8 / 14

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号