TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Apr 15 – Apr 21, 2024
本周最热91

Learn Your Reference Model for Real Good Alignment

Alexey Gorbatovski, Boris Shaposhnikov, Alexey Malakhov +5 authors

A new method, Trust Region DPO (TR-DPO), is proposed to improve policy alignment in reinforcement learning, outperforming Direct Preference Optimization (DPO) by up to 19% on key datasets by updating the reference policy during training.

Reinforcement Learning From Human Feedback (RLHF)Kullback-Leibler divergencereward maximizationSFT policyHF ↗arXiv ↗

33 篇论文 · 按点赞排序

04

TransformerFAM: Feedback attention is working memory

Dongseong Hwang, Weiran Wang, Zhuoyuan Huo +2 authors

Feedback Attention Memory (FAM) enhances Transformer architecture by enabling long-context processing without additional weights, significantly improving performance on large sequences across various model sizes.

43TransformersFeedback Attention MemoryHF ↗arXiv ↗
06

Dynamic Typography: Bringing Words to Life

Zichen Liu, Yihao Meng, Hao Ouyang +4 authors

Dynamic Typography generates coherent and semantically meaningful text animations using neural displacement fields, end-to-end optimization, and perceptual loss regularization, outperforming baseline methods.

39neural displacement fieldsend-to-end optimizationHF ↗arXiv ↗
07

Pre-training Small Base LMs with Fewer Tokens

Sunny Sanyal, Sujay Sanghavi, Alexandros G. Dimakis

Inheritune leverages transformer blocks from large language models and minimal data to create smaller, efficient base models with performance competitive to larger models.

36transformer blocksInherituneHF ↗arXiv ↗
10

COCONut: Modernizing COCO Segmentation

Xueqing Deng, Qihang Yu, Peng Wang +2 authors

COCONut, a reevaluated and expanded version of the COCO dataset, provides high-quality annotations for semantic, instance, and panoptic segmentation tasks, serving as the first large-scale universal segmentation benchmark.

30HF ↗arXiv ↗
12

Compression Represents Intelligence Linearly

Yuzhen Huang, Jinghan Zhang, Zifei Shan +1 authors

Research shows that the ability of large language models to compress text is linearly associated with their performance on various benchmarks related to intelligence, suggesting compression efficiency as a reliable evaluation metric.

28large language modelscompressionHF ↗arXiv ↗
15

Long-form music generation with latent diffusion

Zach Evans, Julian D. Parker, CJ Carr +3 authors

A diffusion-transformer trained on long temporal contexts can generate full-length music tracks with coherent structure at a latent rate of 21.5Hz.

27diffusion-transformercontinuous latent representationHF ↗arXiv ↗
16

EdgeFusion: On-Device Text-to-Image Generation

Thibault Castells, Hyoung-Kyu Song, Tairen Piao +6 authors

Researchers achieve efficient text-to-image generation with reduced sampling steps and architectural optimizations, specifically using a compact Stable Diffusion variant, high-quality datasets, and an advanced distillation process for rapid and accurate image generation on edge devices.

22Stable Diffusiontext-to-image generationHF ↗arXiv ↗
23

Introducing v0.5 of the AI Safety Benchmark from MLCommons

Bertie Vidgen, Adarsh Agrawal, Ahmed M. Ahmed +94 authors

The AI Safety Benchmark v0.5, developed by MLCommons, assesses the safety risks of chat-tuned language models using a taxonomy of 13 hazard categories and includes 43,090 test items.

13chat-tuned language modelshazard categoriesHF ↗arXiv ↗
24

AniClipart: Clipart Animation with Text-to-Video Priors

Ronghuan Wu, Wanchao Su, Kede Ma +1 authors

AniClipart system transforms static clipart images into high-quality motion sequences using Bézier curves, text-to-video priors, and Video Score Distillation Sampling loss for coherent and identity-preserved animations.

13Bézier curvesmotion regularizationHF ↗arXiv ↗
25

On Speculative Decoding for Multimodal Large Language Models

Mukul Gagrani, Raghavv Goel, Wonseok Jeon +3 authors

Speculative decoding enhances the inference efficiency of Multimodal Large Language Models (MLLMs), specifically LLaVA 7B, by using a language-only draft model and a compact LLaVA model with an image adapter.

13Multimodal Large Language ModelsMLLMsHF ↗arXiv ↗
30

Dataset Reset Policy Optimization for RLHF

Jonathan D. Chang, Wenhao Shan, Owen Oertell +4 authors

The DR-PO algorithm leverages reset techniques in RLHF to improve policy optimization by incorporating an offline preference dataset, resulting in better generative performance compared to PPO and DPO.

10Reinforcement LearningRLHFHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号