TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

50 篇论文 · 按点赞排序

33

HallusionBench: You See What You Think? Or You Think What You See? An Image-Context Reasoning Benchmark Challenging for GPT-4V(ision), LLaVA-1.5, and Other Multi-modality Models

Fuxiao Liu, Tianrui Guan, Zongxia Li +4 authors

HallusionBench is a benchmark that highlights language hallucination and visual illusion issues in vision-language models (VLMs), showcasing the limitations of current state-of-the-art models like GPT-4V and LLaVA-1.5.

27Large language modelsvision modelsHF ↗arXiv ↗
35

Contrastive Prefence Learning: Learning from Human Feedback without RL

Joey Hejna, Rafael Rafailov, Harshit Sikchi +4 authors

A new regret-based algorithm, Contrastive Preference Learning (CPL), learns optimal policies directly from human preferences without learning a reward function, addressing optimization challenges in Reinforcement Learning from Human Feedback (RLHF).

25Reinforcement Learning from Human FeedbackRLHFHF ↗arXiv ↗
40

An Early Evaluation of GPT-4V(ision)

Yang Wu, Shilong Wang, Hao Yang +4 authors

GPT-4V demonstrates strong visual understanding but has limitations in language comprehension, handling sensitive data, modalities like depth and audio, and fine visual nuances.

22GPT-4Vvisual understandingHF ↗arXiv ↗
43

Aligning Text-to-Image Diffusion Models with Reward Backpropagation

Mihir Prabhudesai, Anirudh Goyal, Deepak Pathak +1 authors

AlignProp refines text-to-image diffusion models using backpropagation through the denoising process, leveraging low-rank adapters and gradient checkpointing to optimize for various objectives with higher efficiency.

22text-to-image diffusion modelsreinforcement learningHF ↗arXiv ↗
44

ConvNets Match Vision Transformers at Scale

Samuel L. Smith, Andrew Brock, Leonard Berrada +1 authors

ConvNets pre-trained on a large dataset match the performance of Vision Transformers on ImageNet with comparable computational resources.

21ConvNetsVision TransformersHF ↗arXiv ↗
50

Conditional Diffusion Distillation

Kangfu Mei, Mauricio Delbracio, Hossein Talebi +3 authors

A novel single-stage distillation method for generative diffusion models reduces sampling time while maintaining performance across tasks like super-resolution and image editing.

19generative diffusion modelstext-to-image generationHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号