TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jun 12 – Jun 18, 2023

50 篇论文 · 按点赞排序

31

Anticipatory Music Transformer

John Thickstun, David Hall, Chris Donahue +1 authors

A method for anticipatory generation in music control tasks, using interleaved event and control sequences, matches the performance of autoregressive models and produces high-quality accompaniments.

10temporal point processevent processHF ↗arXiv ↗
34

Transformers learn through gradual rank increase

Enric Boix-Adsera, Etai Littwin, Emmanuel Abbe +2 authors

The study examines incremental learning in transformers, proving that the difference in rank between trained and initial weights increases progressively.

10incremental learningtransformersHF ↗arXiv ↗
37

Agile Catching with Whole-Body MPC and Blackbox Policy Learning

Saminda Abeyruwan, Alex Bewley, Nicholas M. Boffi +9 authors

Two methodologies, Model Predictive Control with constrained trajectory optimization and Reinforcement Learning with zeroth-order optimization, are compared for high-speed object catching in robotics, highlighting performance metrics and potential integration strategies.

9Model Predictive Controlconstrained trajectory optimizationHF ↗arXiv ↗
38

Retrieval-Enhanced Contrastive Vision-Text Models

Ahmet Iscen, Mathilde Caron, Alireza Fathi +1 authors

RECO, a retrieval-enhanced technique using a lightweight fusion transformer, improves CLIP's performance on fine-grained tasks by refining embeddings with cross-modal information from external memory.

8contrastive image-text modelsCLIPHF ↗arXiv ↗
39

LOVM: Language-Only Vision Model Selection

Orr Zohar, Shih-Cheng Huang, Kuan-Chieh Wang +1 authors

A new benchmark is introduced for efficient zero-shot evaluation of pre-trained vision-language models without requiring a downstream dataset, focusing on model selection based on text descriptions.

7multi-modal vision-language modelsVLMsHF ↗arXiv ↗
40

SayTap: Language to Quadrupedal Locomotion

Yujin Tang, Wenhao Yu, Jie Tan +3 authors

A novel approach using-foot contact patterns interfaces natural language commands with LLMs to control quadrupedal robots, achieving high success rates and versatility in locomotion tasks.

7large language models (LLMs)foot contact patternsHF ↗arXiv ↗
42

VidEdit: Zero-Shot and Spatially Aware Text-Driven Video Editing

Paul Couairon, Clément Rambour, Jean-Emmanuel Haugeard +1 authors

VidEdit is a zero-shot text-based video editing method combining atlas-based and pre-trained diffusion models with panoptic segmenters and edge detectors to achieve strong temporal and spatial consistency.

6diffusion-based generative modelstext-to-image diffusion modelsHF ↗arXiv ↗
43

Neural Scene Chronology

Haotong Lin, Qianqian Wang, Ruojin Cai +4 authors

A novel scene representation with temporal step function encoding enables reconstruction of time-varying 3D models from internet photos, achieving state-of-the-art view synthesis with independent control over viewpoint, time, and illumination.

6space-time radiance fieldillumination embeddingHF ↗arXiv ↗
45

arXiVeri: Automatic table verification with GPT

Gyungin Shin, Weidi Xie, Samuel Albanie

The paper introduces a task called automatic table verification (AutoTV) to ensure the accuracy of numerical data in tables through source cross-referencing, proposing a benchmark called arXiVeri and evaluating table and cell matching using LLMs.

6automatic table verificationAutoTVHF ↗arXiv ↗
46

Multi-Modal Classifiers for Open-Vocabulary Object Detection

Prannay Kaul, Weidi Xie, Andrew Zisserman

A novel open-vocabulary object detection method uses language descriptions and image exemplars to detect objects unseen during training, achieving superior performance in multi-modal classification.

6open-vocabulary object detectiontwo-stage object detectorHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号