Dynamic Scaling of Unit Tests for Code Reward Modeling
Zeyao Ma, Xiaokang Zhang, Jing Zhang +3 authors
Generating and dynamically scaling unit tests improves the performance of large language models on complex reasoning tasks like code generation.
Explore · 每周精选
发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。
48 篇论文 · 按点赞排序
Zeyao Ma, Xiaokang Zhang, Jing Zhang +3 authors
Generating and dynamically scaling unit tests improves the performance of large language models on complex reasoning tasks like code generation.
Haoran Wei, Youyang Yin, Yumeng Li +5 authors
A new slow-down perception approach for large vision language models aims to enhance their ability in visual reasoning, particularly in accurately copying and understanding geometric figures by breaking down perception into gradual, human-like stages.
Jiawei Lin, Shizhao Sun, Danqing Huang +3 authors
LaDeCo, a novel approach, uses layered design principles in large multimodal models to perform automatic graphic design composition by breaking down tasks into manageable steps.
Jianyi Wang, Zhijie Lin, Meng Wei +4 authors
SeedVR, a diffusion transformer with shifted window attention, effectively handles real-world video restoration with high performance on synthetic and real-world benchmarks.
Marta Skreta, Lazar Atanackovic, Avishek Joey Bose +2 authors
SuperDiff framework combines multiple pre-trained diffusion models efficiently during inference without re-training, enabling diverse image generation and protein structure design.
Tao Wu, Yong Zhang, Xiaodong Cun +6 authors
A novel framework leverages Video Diffusion Models for zero-shot customized video generation by directly using VDM's intrinsic feature extraction and spatial self-attention for feature injection.
Zhaojian Yu, Yilun Zhao, Arman Cohan +1 authors
Evaluates LLMs on self-invoking code generation tasks, highlighting their decline in performance compared to traditional benchmarks.
Or Patashnik, Rinon Gal, Daniil Ostashev +3 authors
Nested Attention enhances image-text alignment by injecting expressive subject representations into cross-attention layers, enabling high identity preservation and multi-domain subject combination.
Yang Li, Dong Du, Linfeng Song +4 authors
HunyuanProver, a language model fine-tuned for interactive theorem proving with LEAN4, achieves state-of-the-art performance on benchmarks by using a scalable data synthesis framework and guided tree search algorithms.
Shaojin Wu, Fei Ding, Mengqi Huang +2 authors
A Cross-Attention Value Mixing Control (VMix) adapter enhances the aesthetic quality of images generated by diffusion models without retraining.
Hua Farn, Hsuan Su, Shachi H Kumar +3 authors
Merging weights of pre- and post-fine-tuned safety-aligned language models improves downstream task performance while preserving safety without additional safety data.
Mahir Labib Dihan, Mohammed Eunus Ali, Md Rizwan Parvez
MapQaTor is a web application that facilitates the creation of map-based QA datasets using diverse maps APIs, improving the efficiency and reliability of geospatial data annotation and evaluation.
Yanlin Feng, Simone Papicchio, Sajjadur Rahman
Property graph views and Cypher queries enable efficient knowledge retrieval from large RDF graphs for LLMs, improving performance on tasks like question answering.
Peihao Wang, Ruisi Cai, Yuehao Wang +4 authors
Polarizing state transition matrices in Structured State Space Models addresses recency bias and over-smoothing, improving long-range token recall and enabling deeper architectures.
Yang Li, Han Meng, Zhenyu Bi +2 authors
PaD-TS is a population-aware diffusion model for time series data generation that better preserves dataset-wide properties while maintaining individual data point authenticity.
Yongle Huang, Haodong Chen, Zhenbang Xu +3 authors
SeFAR, a semi-supervised learning framework using dual-level temporal elements and adaptive regulation, achieves state-of-the-art performance in fine-grained action recognition and enhances multimodal foundation models' understanding of fine-grained semantics.
Risa Shinoda, Kuniaki Saito, Shohei Tanaka +2 authors
SBSFigures is a pre-training dataset for figure QA that generates synthetic chart figures with accurate annotations through a stage-by-stage pipeline, reducing manual effort and code errors.
Jiajun Zhu, Peihao Wang, Ruisi Cai +3 authors
A novel contextualized equivariant position embedding framework, TAPE, enhances positional encodings in transformers, improving robustness and performance across language modeling, arithmetic reasoning, and long-context retrieval tasks.
北京市昌平区探索星信息技术及软件开发工作室
京ICP备2026059466号