Neural Network Diffusion
Kai Wang, Zhaopan Xu, Yukun Zhou +4 authors
Diffusion models can generate high-performing neural network parameters using an autoencoder, producing new subsets of network parameters with comparable or improved performance.
Explore · 每周精选
发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。
Yiran Ding, Li Lyna Zhang, Chengruidong Zhang +5 authors
LongRoPE extends pre-trained LLMs' context window to 2048k tokens with minimal fine-tuning costs and maintains original performance.
50 篇论文 · 按点赞排序
Kai Wang, Zhaopan Xu, Yukun Zhou +4 authors
Diffusion models can generate high-performing neural network parameters using an autoencoder, producing new subsets of network parameters with comparable or improved performance.
Tianyu Zheng, Ge Zhang, Tianhao Shen +5 authors
OpenCodeInterpreter, an open-source system for generating, executing, and refining code, achieves high performance on benchmarks through execution and human feedback, reducing the gap with proprietary systems like GPT-4 Code Interpreter.
Gagan Bhatia, El Moatez Billah Nagoudi, Hasan Cavusoglu +1 authors
FinTral, a multimodal LLM enhanced through domain-specific pretraining, instruction fine-tuning, and RLAIF, outperforms ChatGPT-3.5 and GPT-4 in financial analysis tasks with exceptional zero-shot performance.
Yaroslav Aksenov, Nikita Balagansky, Sofia Maria Lo Cicero Vaina +3 authors
A modification to the Based model kernel enhances its in-context learning capabilities, outperforming existing subquadratic architectures in language modeling tasks on the Pile dataset.
Haoran Li, Qingxiu Dong, Zhengyang Tang +17 authors
GLAN, a general method for instruction tuning of LLMs, uses a taxonomy of human knowledge to generate synthetic instruction data, achieving strong performance across diverse tasks without task-specific training data.
Chien-Yao Wang, I-Hau Yeh, Hong-Yuan Mark Liao
The paper addresses data loss in deep networks through programmable gradient information (PGI) and introduces the Generalized Efficient Layer Aggregation Network (GELAN) to improve parameter utilization and achieve competitive results in object detection.
Zeyu Lu, Zidong Wang, Di Huang +4 authors
The Flexible Vision Transformer adapts to varied image resolutions and aspect ratios through dynamic tokenization and extrapolation techniques, outperforming traditional methods.
Lucas Lehnert, Sainbayar Sukhbaatar, Paul Mcvay +2 authors
Searchformer, a Transformer model, outperforms traditional $A^*$ search by optimizing Sokoban puzzle solutions with fewer steps and demonstrates superior performance on maze navigation and more complex tasks.
Jun Zhan, Junqi Dai, Jiasheng Ye +13 authors
AnyGPT is a multimodal language model using discrete representations to process and generate content across various modalities with performance on par with specialized models.
Nikhil Bhendawade, Irina Belousova, Qichen Fu +3 authors
Speculative Streaming is a method that incorporates drafting into a target language model, enhancing decoding speed without quality loss while using fewer parameters.
Yuri Kuratov, Aydar Bulatov, Petr Anokhin +3 authors
Recursively augmented GPT-2 fine-tuning significantly extends the processing capability of generative models to handle long document sequences beyond $10^6$ elements.
Long Zhao, Nitesh B. Gundavarapu, Liangzhe Yuan +16 authors
VideoPrism, a pretrained video encoder, achieves top performance across various video understanding tasks by utilizing global-local distillation and token shuffling of semantic video embeddings enhanced with associated text.
Chiyu Zhang, Yifei Sun, Jun Chen +7 authors
The SPAR framework enhances content recommendations by using pretrained language models, poly-attention layers, and large language models to effectively process long user engagement histories and predict user-item interactions.
Zhaoyang Lv, Nickolas Charron, Pierre Moulon +21 authors
The AEA Dataset includes multimodal sensor and machine perception data from daily activities, enabling applications like neural scene reconstruction and prompted segmentation.
Ajay Patel, Colin Raffel, Chris Callison-Burch
DataDreamer is an open-source Python library facilitating LLM workflows and promoting open science and reproducibility in NLP research.
Shanchuan Lin, Anran Wang, Xiao Yang
A diffusion distillation method using progressive and adversarial techniques achieves high-quality one-step/few-step text-to-image generation with improved mode coverage and is available as a distilled model.
Md Mohaiminul Islam, Ngan Ho, Xitong Yang +3 authors
Video ReCap is a recursive model for video captioning that handles videos of varying lengths and outputs captions at multiple hierarchical levels using curriculum learning.
Bryan Wang, Yuliang Li, Zhaoyang Lv +3 authors
Video creation has become increasingly popular, yet the expertise and effort required for editing often pose barriers to beginners. In this paper, we explore the integration of large language models (LLMs) into the video editing workflow to reduce these barriers. Our design vision is embodied in LAVE, a novel system that provides LLM-powered agent assistance and language-augmented editing features. LAVE automatically generates language descriptions for the user's footage, serving as the foundation for enabling the LLM to process videos and assist in editing tasks. When the user provides editing objectives, the agent plans and executes relevant actions to fulfill them. Moreover, LAVE allows users to edit videos through either the agent or direct UI manipulation, providing flexibility and enabling manual refinement of agent actions. Our user study, which included eight participants ranging from novices to proficient editors, demonstrated LAVE's effectiveness. The results also shed light on user perceptions of the proposed LLM-assisted editing paradigm and its impact on users' creativity and sense of co-creation. Based on these findings, we propose design implications to inform the future development of agent-assisted content editing.
Zhengbao Jiang, Zhiqing Sun, Weijia Shi +6 authors
Pre-instruction-tuning enhances large language models' ability to learn from documents by first exposing them to question-answer pairs.
Muhammad Maaz, Hanoona Rasheed, Abdelrahman Shaker +6 authors
A multilingual multimodal model called Palo enhances visual reasoning across 10 languages using a semi-automated translation approach and achieves substantial performance improvements across multiple scales.
Qianqian Xie, Weiguang Han, Zhengyu Chen +31 authors
FinBen is an open-sourced financial evaluation benchmark for LLMs, assessing their capabilities across various tasks and difficulty levels, revealing strengths and weaknesses for targeted enhancements.
Yuzhuang Xu, Xu Han, Zonghan Yang +5 authors
OneBit, a 1-bit quantization-aware training framework, allows LLMs to achieve robust performance with 1-bit weight matrices, reducing storage and computational overhead.
Jacky Liang, Fei Xia, Wenhao Yu +47 authors
Language Model Predictive Control (LMPC) improves the teachability of robot code-writing LLMs through fine-tuning, enhancing task success rates and reducing human corrections.
Minsuk Kahng, Ian Tenney, Mahima Pushkarna +7 authors
LLM Comparator is a visual analytics tool for interactive evaluation of large language models' performance compared to baselines, addressing scalability and interpretability challenges.
Lin Ning, Luyang Liu, Jiaxing Wu +6 authors
User-LLM framework uses user embeddings and Perceiver layers to improve LLM performance on long sequence tasks and user understanding efficiently.
Baichuan Zhou, Ying Hu, Xi Weng +5 authors
TinyLLaVA framework demonstrates that small-scale Large Multimodal Models can achieve comparable performance to larger models through improved data quality and training recipes.
Willi Menapace, Aliaksandr Siarohin, Ivan Skorokhodov +8 authors
Snap Video, a transformer-based model, addresses video generation challenges by extending EDM for spatial and temporal redundancy, achieving superior quality, consistency, and speed compared to existing U-Net based methods.
Byung-Kwan Lee, Beomchan Park, Chae Won Kim +1 authors
The study proposes CoLLaVO, a Visual Language Model enhanced with crayon prompt tuning and Dual QLoRA to improve object-level image understanding and zero-shot performance in Vision Language tasks.
Johan Obando-Ceron, Aaron Courville, Pablo Samuel Castro
Gradual magnitude pruning in deep reinforcement learning enhances parameter effectiveness, improving performance and achieving scaling law efficiency.
北京市昌平区探索星信息技术及软件开发工作室
京ICP备2026059466号