TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Feb 19 – Feb 25, 2024
本周最热116

LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens

Yiran Ding, Li Lyna Zhang, Chengruidong Zhang +5 authors

LongRoPE extends pre-trained LLMs' context window to 2048k tokens with minimal fine-tuning costs and maintains original performance.

Large language modelsfine-tuning costscontext windowLongRoPEHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

Neural Network Diffusion

Kai Wang, Zhaopan Xu, Yukun Zhou +4 authors

Diffusion models can generate high-performing neural network parameters using an autoencoder, producing new subsets of network parameters with comparable or improved performance.

101diffusion modelsautoencoderHF ↗arXiv ↗
08

FiT: Flexible Vision Transformer for Diffusion Model

Zeyu Lu, Zidong Wang, Di Huang +4 authors

The Flexible Vision Transformer adapts to varied image resolutions and aspect ratios through dynamic tokenization and extrapolation techniques, outperforming traditional methods.

48diffusion modelsDiffusion TransformersHF ↗arXiv ↗
13

VideoPrism: A Foundational Visual Encoder for Video Understanding

Long Zhao, Nitesh B. Gundavarapu, Liangzhe Yuan +16 authors

VideoPrism, a pretrained video encoder, achieves top performance across various video understanding tasks by utilizing global-local distillation and token shuffling of semantic video embeddings enhanced with associated text.

41VideoPrismmasked autoencodingHF ↗arXiv ↗
15

Aria Everyday Activities Dataset

Zhaoyang Lv, Nickolas Charron, Pierre Moulon +21 authors

The AEA Dataset includes multimodal sensor and machine perception data from daily activities, enabling applications like neural scene reconstruction and prompted segmentation.

31egocentric multimodal datasetneural scene reconstructionHF ↗arXiv ↗
18

Video ReCap: Recursive Captioning of Hour-Long Videos

Md Mohaiminul Islam, Ngan Ho, Xitong Yang +3 authors

Video ReCap is a recursive model for video captioning that handles videos of varying lengths and outputs captions at multiple hierarchical levels using curriculum learning.

27recursive video captioningvideo hierarchiesHF ↗arXiv ↗
19

LAVE: LLM-Powered Agent Assistance and Language Augmentation for Video Editing

Bryan Wang, Yuliang Li, Zhaoyang Lv +3 authors

Video creation has become increasingly popular, yet the expertise and effort required for editing often pose barriers to beginners. In this paper, we explore the integration of large language models (LLMs) into the video editing workflow to reduce these barriers. Our design vision is embodied in LAVE, a novel system that provides LLM-powered agent assistance and language-augmented editing features. LAVE automatically generates language descriptions for the user's footage, serving as the foundation for enabling the LLM to process videos and assist in editing tasks. When the user provides editing objectives, the agent plans and executes relevant actions to fulfill them. Moreover, LAVE allows users to edit videos through either the agent or direct UI manipulation, providing flexibility and enabling manual refinement of agent actions. Our user study, which included eight participants ranging from novices to proficient editors, demonstrated LAVE's effectiveness. The results also shed light on user perceptions of the proposed LLM-assisted editing paradigm and its impact on users' creativity and sense of co-creation. Based on these findings, we propose design implications to inform the future development of agent-assisted content editing.

27HF ↗arXiv ↗
21

PALO: A Polyglot Large Multimodal Model for 5B People

Muhammad Maaz, Hanoona Rasheed, Abdelrahman Shaker +6 authors

A multilingual multimodal model called Palo enhances visual reasoning across 10 languages using a semi-automated translation approach and achieves substantial performance improvements across multiple scales.

24Large Multilingual Multimodal Modelvisual reasoningHF ↗arXiv ↗
29

CoLLaVO: Crayon Large Language and Vision mOdel

Byung-Kwan Lee, Beomchan Park, Chae Won Kim +1 authors

The study proposes CoLLaVO, a Visual Language Model enhanced with crayon prompt tuning and Dual QLoRA to improve object-level image understanding and zero-shot performance in Vision Language tasks.

21Large Language Models (LLMs)Vision Language Models (VLMs)HF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号