iFormer: Integrating ConvNet and Transformer for Mobile Application
Chuanyang Zheng
iFormer combines convolution and self-attention to create a lightweight, high-performance mobile vision network with low latency across various tasks.
Explore · 每周精选
发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。
50 篇论文 · 按点赞排序
Chuanyang Zheng
iFormer combines convolution and self-attention to create a lightweight, high-performance mobile vision network with low latency across various tasks.
J. Pablo Muñoz, Jinjie Yuan, Nilesh Jain
Low-rank adapters and NAS techniques are combined to create efficient, parameter-reduced Large Language Models suitable for resource-limited environments.
Samira Abnar, Harshay Shah, Dan Busbridge +3 authors
Research on sparse Mixture-of-Experts (MoEs) shows an optimal level of sparsity improves both training efficiency and model performance, providing insights for designing efficient architectures.
Tiansheng Huang, Sihao Hu, Fatih Ilhan +2 authors
A new attack method, Virus, bypasses guardrail filters by slightly modifying harmful data, demonstrating the unreliability of guardrails in preventing harmful fine-tuning on pre-trained language models.
Ryan Ehrlich, Bradley Brown, Jordan Juravsky +3 authors
CodeMonkeys iteratively generates and tests code edits to resolve real-world GitHub issues, scaling test-time compute through multiple iterations and parallel trajectories, and outperforms individual models by combining candidate solutions.
Makoto Shing, Kou Misaki, Han Bao +2 authors
Temporally Adaptive Interpolated Distillation (TAID) addresses capacity gap, mode averaging, and mode collapse between large and small models, enhancing knowledge distillation and leading to efficient compact foundation models for language and vision-language tasks.
Weixin Liang, Junhong Shen, Genghan Zhang +3 authors
Mixture-of-Mamba extends modality-aware sparsity to State Space Models, improving multi-modal pretraining efficiency across text, image, and speech.
Sara Kothari, Ayush Gupta
A fine-tuned LLM approach for semantic QA over EHRs using FHIR resources outperforms GPT-4 in both task performance and model size, with evaluation techniques including sequential fine-tuning and self-evaluation.
Shaofei Wang, Tomas Simon, Igor Santesteban +15 authors
A new approach for relightable full-body avatars uses zonal harmonics for local light transport and a shadow network for non-local effects, enabling superior generalization under novel lighting and poses.
Sankalp KJ, Ashutosh Kumar, Laxmaan Balaji +4 authors
IndicMMLU-Pro evaluates multilingual large language models across various Indic languages using a comprehensive benchmark that covers comprehension, reasoning, and generation tasks.
Huayu Chen, Kai Jiang, Kaiwen Zheng +3 authors
Guidance-Free Training (GFT) matches the performance of Classifier-Free Guidance (CFG) while reducing computational costs by eliminating guided sampling and allowing for training from scratch.
Paul Gavrikov, Jovita Lukasik, Steffen Jung +5 authors
Vision-language models exhibit biases that can be influenced by language prompts, often showing more shape bias compared to pure vision models.
Akashah Shabbir, Mohammed Zumri, Mohammed Bennamoun +2 authors
GeoPixel, an end-to-end RS-LMM, achieves superior pixel-level comprehension in remote sensing images through fine-grained grounding and interleaved mask generation in conversations.
Faria Huq, Zora Zhiruo Wang, Frank F. Xu +4 authors
CowPilot is a framework that combines human-agent collaboration in web navigation, demonstrating high task success and efficiency by reducing user steps.
Xiaoyang Wang, Hongming Zhang, Tao Ge +3 authors
A large-scale data synthesis approach equips LLMs with character generalization capabilities through supervised fine-tuning, achieving performance comparable to GPT-4 on role-playing dialogue.
Yang You, Yixin Li, Congyue Deng +2 authors
Evaluating and enhancing 3D equivariance in ViT-based models improves performance in tasks like pose estimation, tracking, and semantic transfer through simple fine-tuning with 3D correspondences.
Thibaud Leteno, Irina Proskurina, Antoine Gourru +4 authors
Histoires Morales, a French moral reasoning dataset, addresses gaps in understanding how language models align with human values in French by providing culturally adapted data and annotations.
Mohamed Elfeki, Rui Liu, Chad Voegele
Encoder-decoder architectures offer superior efficiency and performance compared to decoder-only models for small language models and low-resource environments, especially when enhanced with knowledge distillation and modern embeddings.
Juan Ramirez, Ignacio Hounie, Juan Elenter +4 authors
Feasible Learning trains models by ensuring satisfactory performance on each sample, using a primal-dual approach and minimal norm slack variables, improving tail behavior with slight impact on average performance.
Zheng Chong, Wenqing Zhang, Shiyue Zhang +6 authors
CatV2TON, a vision-based virtual try-on method using a diffusion transformer model, achieves high-quality results for both image and video try-on tasks, including efficient long-video generation through overlapping clip-based inference and adaptive clip normalization.
北京市昌平区探索星信息技术及软件开发工作室
京ICP备2026059466号