Neural Network Diffusion
Kai Wang, Zhaopan Xu, Yukun Zhou +4 authors
Diffusion models can generate high-performing neural network parameters using an autoencoder, producing new subsets of network parameters with comparable or improved performance.
Explore · 每周精选
发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。
598 篇论文 · 按点赞排序
Kai Wang, Zhaopan Xu, Yukun Zhou +4 authors
Diffusion models can generate high-performing neural network parameters using an autoencoder, producing new subsets of network parameters with comparable or improved performance.
Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo +22 authors
Emu3, a transformer-based multimodal model trained exclusively with next-token prediction, outperforms existing diffusion and compositional models in generation and perception tasks.
Taiming Lu, Tianmin Shu, Junfei Xiao +8 authors
GenEx generates 3D environments from a single image, enabling AI agents to explore and interact with a consistent, expansive space through guided generative imagination.
Chenglei Si, Yanzhe Zhang, Zhengyuan Yang +2 authors
Evaluates the performance of multimodal LLMs, including GPT-4V and Gemini Pro Vision, on converting visual designs to code using a new Design2Code task and benchmarks.
Jianwen Jiang, Chao Liang, Jiaqi Yang +3 authors
An audio-only conditioned video diffusion model, Loopy, improves natural and high-quality human video generation without auxiliary spatial signals.
Pan Zhang, Xiaoyi Dong, Yuhang Cao +26 authors
A framework combining disentangled perception, reasoning, and memory mechanisms enables real-time interaction for long-term multimodal AI systems.
Peiyuan Zhang, Guangtao Zeng, Tianduo Wang +1 authors
TinyLlama, a compact 1.1B language model, leverages FlashAttention to achieve high performance in downstream tasks with enhanced computational efficiency.
Daixuan Cheng, Yuxian Gu, Shaohan Huang +3 authors
Instruction Pre-Training enhances language models by generating and incorporating instruction-response pairs into unsupervised multitask pre-training.
Yuan Yao, Tianyu Yu, Ao Zhang +20 authors
MiniCPM-V presents a series of efficient Multimodal Large Language Models optimized for end-side deployment, offering high performance and practical usability compared to larger models.
Vadim Titov, Madina Khalmatova, Alexandra Ivanova +2 authors
A modified diffusion sampling process using self-guidance and noise rescaling achieves high-quality image editing with preserved structure and appearance without requiring diffusion model fine-tuning.
Zhenghao Lin, Zhibin Gou, Yeyun Gong +8 authors
Rho-1, a novel language model using Selective Language Modeling, improves efficiency and performance by selectively training on useful tokens rather than all tokens in the corpus.
Shijia Yang, Bohan Zhai, Quanzeng You +3 authors
Correlation between cross-modal alignment and vision representation improves performance in multimodal large language models, enabling identification and training of optimal vision representation with reduced computational cost.
Pan Zhang, Xiaoyi Dong, Yuhang Zang +24 authors
A large-vision language model with long-contextual capabilities supports various comprehension and composition tasks, achieving GPT-4V level performance using 7B parameters.
Shibo Hao, Sainbayar Sukhbaatar, DiJia Su +4 authors
A new reasoning paradigm, Coconut, utilizes continuous representations in latent space to enhance LLM performance on complex reasoning tasks, particularly those requiring backtracking.
Zesen Cheng, Hang Zhang, Kehan Li +6 authors
A tile-based computation strategy optimizes contrastive loss by partitioning calculations, enabling large batch sizes without significant memory or speed sacrifices.
Rawal Khirodkar, Timur Bagautdinov, Julieta Martinez +5 authors
Sapiens, a family of vision models for human-centric tasks, achieves superior performance with self-supervised pretraining and can be easily fine-tuned for 2D pose estimation, body-part segmentation, depth estimation, and surface normal prediction.
Pierre Colombo, Telmo Pessoa Pires, Malik Boudiaf +8 authors
SaulLM-7B, a large language model with 7 billion parameters, excels in legal text comprehension and generation using instructional fine-tuning on a legal corpus.
Zifan Zheng, Yezhaohui Wang, Yuxin Huang +4 authors
This survey explores the interpretability and underlying mechanisms of attention heads within Large Language Models to understand their reasoning processes.
Junnan Liu, Hongwei Liu, Linchen Xiao +6 authors
G-Pass@k and LiveMathBench assess Large Language Models' reasoning capabilities and consistency using complex mathematical problems.
Florian Bordes, Richard Yuanzhe Pang, Anurag Ajay +38 authors
Introduction to vision-language models (VLMs) covering their applications, training, evaluation, and extension to videos, addressing challenges in mapping visual data to language.
Alexey Gorbatovski, Boris Shaposhnikov, Alexey Malakhov +5 authors
A new method, Trust Region DPO (TR-DPO), is proposed to improve policy alignment in reinforcement learning, outperforming Direct Preference Optimization (DPO) by up to 19% on key datasets by updating the reference policy during training.
Jianfeng Xiang, Zelong Lv, Sicheng Xu +6 authors
A 3D generation method using a unified SLAT representation and rectified flow transformers achieves high-quality results across different formats and conditions.
Nicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham +10 authors
A model-stealing attack is introduced that can extract detailed information such as the embedding projection layer from black-box language models with minimal cost.
Dan Biderman, Jose Gonzalez Ortiz, Jacob Portes +9 authors
LoRA, a parameter-efficient finetuning method for large language models, underperforms full finetuning in target domains but provides better regularization and maintains diverse generation compared to other techniques.
Kevin Qinghong Lin, Linjie Li, Difei Gao +6 authors
ShowUI is a vision-language-action model that enhances GUI assistants by using UI-guided token selection and interleaved vision-language-action streaming, achieving high accuracy and efficiency in zero-shot screenshot grounding across different environments.
Alexander Nikulin, Ilya Zisman, Alexey Zemtsov +3 authors
XLand-100B is a large-scale dataset for in-context reinforcement learning, containing extensive learning histories in the XLand-MiniGrid environment.
Philippe Laban, Alexander R. Fabbri, Caiming Xiong +1 authors
The SummHay task evaluates LLMs and RAG systems on summarizing long-context documents by identifying relevant insights and citing sources, revealing challenges even for systems with document relevance signals.
Min Shi, Fuxiao Liu, Shihao Wang +12 authors
Mixture of vision encoders and resolutions in multimodal large language models improves performance through concatenation of visual tokens and a Pre-Alignment mechanism, leading to superior results on benchmarks.
OpenAI, Aaron Hurst, Adam Lerer +416 authors
GPT-4o is an omnimodal autoregressive model trained to handle text, audio, image, and video inputs, offering high-performance outputs across these modalities, with particular strengths in vision and audio.
Yadong Li, Haoze Sun, Mingan Lin +24 authors
Baichuan-Omni, a 7B open-source Multimodal Large Language Model, excels in processing image, video, audio, and text, showcasing competitive performance across multimodal benchmarks.
北京市昌平区探索星信息技术及软件开发工作室
京ICP备2026059466号