OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
Zhiyong Wu, Zhenyu Wu, Fangzhi Xu +8 authors
OS-Atlas improves GUI agent performance through a large open-source GUI grounding dataset and model innovations.
Explore · 每周精选
发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。
50 篇论文 · 按点赞排序
Zhiyong Wu, Zhenyu Wu, Fangzhi Xu +8 authors
OS-Atlas improves GUI agent performance through a large open-source GUI grounding dataset and model innovations.
Erik Wijmans, Brody Huval, Alexander Hertzberg +2 authors
Cut Cross-Entropy reduces the memory footprint of large language models during training by efficiently computing cross-entropy loss without materializing logits for all tokens.
Yifan Xu, Xiao Liu, Xueqiao Sun +7 authors
AndroidLab provides a systematic framework for training and evaluating Android agents, supporting both large language models and multimodal models, and improves their task success rates.
Enrico Fini, Mustafa Shukor, Xiujun Li +13 authors
AIMV2, a multimodal vision encoder paired with an autoregressive decoder, achieves superior performance in vision benchmarks and multimodal image understanding tasks compared to contrastive models.
Xudong Lu, Yinghao Chen, Cheng Chen +19 authors
BlueLM-V-3B optimizes the deployment of multimodal large language models on mobile platforms through redesigned dynamic resolution and hardware-aware system optimizations, achieving high performance with a small model size and fast generation speed.
Zhen Huang, Haoyang Zou, Xuefeng Li +7 authors
Knowledge distillation from O1 model combined with fine-tuning achieves superior performance on complex mathematical reasoning and generalizes to open-ended QA tasks, promoting transparency in AI research.
Yew Ken Chia, Liying Cheng, Hou Pong Chan +5 authors
A benchmark and framework named M-LongDoc evaluates and improves large multimodal models for document reading, focusing on open-ended solutions and retrieval tasks over lengthy documents.
Di Zhang, Jingdi Lei, Junxian Li +10 authors
The Critic-V framework leverages an Actor-Critic paradigm to enhance multimodal reasoning in vision-language models by decoupling reasoning and critique processes, significantly improving performance and reliability.
Jooyoung Choi, Chaehun Shin, Yeongtak Oh +2 authors
The Style-friendly SNR sampler modifies the noise level distribution during fine-tuning to improve style alignment in diffusion models, enabling better capture of unique artistic styles.
Dawei Li, Bohan Jiang, Liangjie Huang +10 authors
A survey of LLM-based judgment and assessment explores definitions, taxonomy, benchmarks, challenges, and future directions for the LLM-as-a-judge paradigm.
Chaehun Shin, Jooyoung Choi, Heeseung Kim +1 authors
Diptych Prompting is a zero-shot text-to-image generation method that enhances subject alignment and attention between panels during inpainting, resulting in visually preferable images across various applications.
Tianbin Li, Yanzhou Su, Wei Li +15 authors
GMAI-VL, a vision-language model trained on the GMAI-VL-5.5M dataset, achieves top performance in multimodal medical tasks by integrating visual and textual information.
Zhangchen Xu, Fengqing Jiang, Luyao Niu +2 authors
The Larger Models' Paradox reveals that larger models are not always better teachers for fine-tuning smaller models, and a new metric, Compatibility-Adjusted Reward (CAR), is introduced to measure and improve the effectiveness of response generators.
Weiquan Huang, Aoqi Wu, Yifan Yang +8 authors
LLM2CLIP leverages large language models to enhance CLIP's multimodal representation learning, enabling better handling of complex image captions by fine-tuning the LLM and using contrastive learning.
Shenghai Yuan, Jinfa Huang, Xianyi He +5 authors
ConsisID is a tuning-free, DiT-based IPT2V model that uses frequency-aware identity control to generate high-quality, identity-preserving videos by decomposing facial features into low and high-frequency elements.
Dang Nguyen, Viet Dac Lai, Seunghyun Yoon +9 authors
A new LLM agent framework enables dynamic action creation and reuse by generating and executing programs, offering greater flexibility and improved performance in complex, real-world scenarios.
Noam Rotstein, Gal Yona, Daniel Silver +3 authors
Image editing is improved using pretrained video models to ensure accurate edits and preserve image fidelity by reformulating the process as a smooth temporal transition.
Haolin Chen, Yihao Feng, Zuxin Liu +8 authors
LaTent Reasoning Optimization (LaTRO) enhances LLMs' reasoning capabilities through variational optimization of latent distributions, improving zero-shot accuracy on complex reasoning tasks.
Zehan Qi, Xiao Liu, Iat Long Iong +10 authors
WebRL enhances open LLMs as web agents through a self-evolving curriculum, robust outcome-supervised reward model, and adaptive reinforcement learning strategies, achieving high success rates compared to proprietary models.
Akari Asai, Jacqueline He, Rulin Shao +22 authors
OpenScholar, a specialized language model, enhances scientific query synthesis by retrieving relevant passages and providing citation-backed answers, outperforming existing models in correctness and citation accuracy.
北京市昌平区探索星信息技术及软件开发工作室
京ICP备2026059466号