TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jun 3 – Jun 9, 2024
本周最热75

ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Lin Chen, Xilin Wei, Jinsong Li +12 authors

The ShareGPT4Video series enhances video understanding and generation using dense captions and an efficient captioning strategy by addressing temporal and spatial challenges in video annotation.

GPT4VLVLMsT2VMsdense captionsHF ↗arXiv ↗

41 篇论文 · 按点赞排序

04

Block Transformer: Global-to-Local Language Modeling for Fast Inference

Namgyu Ho, Sangmin Bae, Taehyeon Kim +6 authors

The Block Transformer architecture enhances inference throughput by applying global-to-local modeling to autoregressive transformers, reducing inference bottlenecks through hierarchical processing and block-level self-attention.

39Block Transformerhierarchical global-to-local modelingHF ↗arXiv ↗
05

Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Philip Anastassiou, Jiawei Chen, Jitong Chen +43 authors

Seed-TTS is a family of large-scale TTS models that generate high-quality speech with in-context learning, superior controllability, and a non-autoregressive variant using diffusion-based architecture that does not rely on pre-estimated phoneme durations.

39autoregressive text-to-speechspeech generationHF ↗arXiv ↗
07

Parrot: Multilingual Visual Instruction Tuning

Hai-Long Sun, Da-Wei Zhou, Yang Li +8 authors

Parrot enhances multimodal language models with multilingual visual token alignment using textual guidance and Mixture-of-Experts, achieving state-of-the-art performance on multilingual benchmarks.

34Multimodal Large Language ModelsMLLMsHF ↗arXiv ↗
08

To Believe or Not to Believe Your LLM

Yasin Abbasi Yadkori, Ilja Kuzborskij, András György +1 authors

Information-theoretic uncertainty quantification in large language models reliably detects epistemic uncertainty, allowing detection of hallucinations in both single- and multi-answer responses.

34large language modelsuncertainty quantificationHF ↗arXiv ↗
13

SF-V: Single Forward Video Generation Model

Zhixing Zhang, Yanyu Li, Yushu Wu +9 authors

Adversarial training transforms a multi-step diffusion video generation model into a single-step high-quality video synthesis model with reduced computational cost.

24diffusion-based video generationiterative denoisingHF ↗arXiv ↗
21

Self-Improving Robust Preference Optimization

Eugene Choi, Arash Ahmadian, Matthieu Geist +2 authors

SRPO, a self-improving offline RLHF framework, achieves robustness to out-of-distribution tasks by optimizing a min-max objective that jointly enhances self-improvement and generative policies, leading to superior performance compared to DPO.

19PPODPOHF ↗arXiv ↗
22

pOps: Photo-Inspired Diffusion Operators

Elad Richardson, Yuval Alaluf, Ali Mahdavi-Amiri +1 authors

pOps framework uses Diffusion Prior model to train semantic operators in CLIP image embedding space for text-guided image generation.

17CLIP image embedding spaceIP-AdapterHF ↗arXiv ↗
25

Guiding a Diffusion Model with a Bad Version of Itself

Tero Karras, Miika Aittala, Tuomas Kynkäänniemi +3 authors

Using smaller, less-trained versions of the model to guide generation leads to better prompt alignment and higher image quality without reducing variation in diffusion models.

16diffusion modelsimage qualityHF ↗arXiv ↗
28

Open-Endedness is Essential for Artificial Superhuman Intelligence

Edward Hughes, Michael Dennis, Jack Parker-Holder +5 authors

In recent years there has been a tremendous surge in the general capabilities of AI systems, mainly fuelled by training foundation models on internetscale data. Nevertheless, the creation of openended, ever self-improving AI remains elusive. In this position paper, we argue that the ingredients are now in place to achieve openendedness in AI systems with respect to a human observer. Furthermore, we claim that such open-endedness is an essential property of any artificial superhuman intelligence (ASI). We begin by providing a concrete formal definition of open-endedness through the lens of novelty and learnability. We then illustrate a path towards ASI via open-ended systems built on top of foundation models, capable of making novel, humanrelevant discoveries. We conclude by examining the safety implications of generally-capable openended AI. We expect that open-ended foundation models will prove to be an increasingly fertile and safety-critical area of research in the near future.

13foundational modelsopen-endednessHF ↗arXiv ↗
30

4Diffusion: Multi-view Video Diffusion Model for 4D Generation

Haiyu Zhang, Xinyuan Chen, Yaohui Wang +3 authors

A novel 4D generation pipeline, 4Diffusion, uses a unified diffusion model with a learnable motion module to generate spatial-temporally consistent 4D content from monocular video, and introduces a 4D-aware Score Distillation Sampling loss and anchor loss to improve performance.

13diffusion generative modelsunified diffusion modelHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号