TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Oct 23 – Oct 29, 2023
本周最热53

A Picture is Worth a Thousand Words: Principled Recaptioning Improves Image Generation

Eyal Segalis, Dani Valevski, Danny Lumen +2 authors

Relabeling a text-to-image dataset with a specialized automatic captioning model improves image quality and semantic alignment in diffusion models.

text-to-image diffusion modelsLAION datasetStable DiffusionFIDHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

Matryoshka Diffusion Models

Jiatao Gu, Shuangfei Zhai, Yizhe Zhang +2 authors

Matryoshka Diffusion Models use a NestedUNet architecture for joint denoising at multiple resolutions, enabling efficient high-resolution image and video synthesis.

46diffusion modelshigh-resolution image and video synthesisHF ↗arXiv ↗
03

In-Context Learning Creates Task Vectors

Roee Hendel, Mor Geva, Amir Globerson

In-Context Learning in Large Language Models can be understood as compressing a training set into a task vector that modulates a transformer for output generation.

43in-context learninglarge language modelsHF ↗arXiv ↗
05

JudgeLM: Fine-tuned Large Language Models are Scalable Judges

Lianghui Zhu, Xinggang Wang, Xinlong Wang

Large Language Models fine-tuned as scalable judges (JudgeLM) achieve state-of-the-art performance in evaluating open-ended benchmarks through a comprehensive dataset and benchmark, enhancing judgment efficiency and accuracy.

35Large Language Modelsfine-tuningHF ↗arXiv ↗
08

HallusionBench: You See What You Think? Or You Think What You See? An Image-Context Reasoning Benchmark Challenging for GPT-4V(ision), LLaVA-1.5, and Other Multi-modality Models

Fuxiao Liu, Tianrui Guan, Zongxia Li +4 authors

HallusionBench is a benchmark that highlights language hallucination and visual illusion issues in vision-language models (VLMs), showcasing the limitations of current state-of-the-art models like GPT-4V and LLaVA-1.5.

27Large language modelsvision modelsHF ↗arXiv ↗
09

Contrastive Prefence Learning: Learning from Human Feedback without RL

Joey Hejna, Rafael Rafailov, Harshit Sikchi +4 authors

A new regret-based algorithm, Contrastive Preference Learning (CPL), learns optimal policies directly from human preferences without learning a reward function, addressing optimization challenges in Reinforcement Learning from Human Feedback (RLHF).

25Reinforcement Learning from Human FeedbackRLHFHF ↗arXiv ↗
12

An Early Evaluation of GPT-4V(ision)

Yang Wu, Shilong Wang, Hao Yang +4 authors

GPT-4V demonstrates strong visual understanding but has limitations in language comprehension, handling sensitive data, modalities like depth and audio, and fine visual nuances.

22GPT-4Vvisual understandingHF ↗arXiv ↗
13

ConvNets Match Vision Transformers at Scale

Samuel L. Smith, Andrew Brock, Leonard Berrada +1 authors

ConvNets pre-trained on a large dataset match the performance of Vision Transformers on ImageNet with comparable computational resources.

21ConvNetsVision TransformersHF ↗arXiv ↗
17

SALMONN: Towards Generic Hearing Abilities for Large Language Models

Changli Tang, Wenyi Yu, Guangzhi Sun +6 authors

SALMONN, an integrated multimodal model combining a pre-trained text-based LLM with speech and audio encoders, demonstrates competitive performance and emergent abilities in various speech and audio tasks.

17pre-trained text-based large language model (LLM)speech encodersHF ↗arXiv ↗
22

Controlled Decoding from Language Models

Sidharth Mudgal, Jong Lee, Harish Ganapathy +10 authors

Controlled decoding is an off-policy reinforcement learning method that uses a prefix scorer to guide language model generation towards high rewards and can handle multiple objectives without additional complexity.

14controlled decodingoff-policy reinforcement learningHF ↗arXiv ↗
28

Detecting Pretraining Data from Large Language Models

Weijia Shi, Anirudh Ajith, Mengzhou Xia +5 authors

Researchers propose Min-K% Prob, a method for detecting pretraining data in large language models without needing knowledge of the pretraining corpus, using a benchmark that supports gold truth detection.

11large language modelspretraining data detectionHF ↗arXiv ↗
30

TiC-CLIP: Continual Training of CLIP Models

Saurabh Garg, Mehrdad Farajtabar, Hadi Pouransari +5 authors

Web-scale Time-Continual (TiC) benchmarks are introduced to evaluate and improve the temporal robustness of vision-language models through efficient continual learning methods.

10time-continual (TiC) benchmarksvision-language modelsHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号