TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Oct 20 – Oct 26, 2025

50 篇论文 · 按点赞排序

31

DyPE: Dynamic Position Extrapolation for Ultra High Resolution Diffusion

Noam Issachar, Guy Yariv, Sagie Benaim +3 authors

Dynamic Position Extrapolation (DyPE) enhances ultra-high-resolution image generation by dynamically adjusting positional encodings in pre-trained diffusion transformers, achieving state-of-the-art fidelity without additional sampling cost.

37Diffusion Transformer modelsself-attention mechanismHF ↗arXiv ↗
33

QueST: Incentivizing LLMs to Generate Difficult Problems

Hanxu Hu, Xingxing Zhang, Jannis Vamvas +2 authors

QueST, a framework combining difficulty-aware graph sampling and fine-tuning, generates large-scale synthetic coding problems to enhance the performance of large language models in competitive coding and reasoning tasks.

36difficulty-aware graph samplingdifficulty-aware rejection fine-tuningHF ↗arXiv ↗
36

AION-1: Omnimodal Foundation Model for Astronomical Sciences

Liam Parker, Francois Lanusse, Jeff Shen +24 authors

AION-1, a family of large-scale multimodal foundation models, integrates diverse astronomical data using tokenization and transformer-based modeling, achieving strong performance across various downstream tasks.

33multimodal foundation modelstokenizationHF ↗arXiv ↗
38

Pico-Banana-400K: A Large-Scale Dataset for Text-Guided Image Editing

Yusu Qian, Eli Bocek-Rivele, Liangchen Song +5 authors

Pico-Banana-400K is a large-scale, high-quality dataset for instruction-based image editing, featuring diverse edit pairs, multi-turn editing, preference subsets, and long-short instruction pairs, enabling comprehensive research and benchmarking.

30multimodal modelstext-guided image editingHF ↗arXiv ↗
41

Chronos-2: From Univariate to Universal Forecasting

Abdul Fatir Ansari, Oleksandr Shchur, Jaris Küken +20 authors

Chronos-2, a pretrained model with a group attention mechanism, achieves state-of-the-art performance in zero-shot univariate, multivariate, and covariate-informed forecasting tasks.

28pretrained modelgroup attention mechanismHF ↗arXiv ↗
42

IF-VidCap: Can Video Caption Models Follow Instructions?

Shihao Li, Yuanxing Zhang, Jiangtao Wu +20 authors

A new benchmark, IF-VidCap, evaluates video captioning models on instruction-following capabilities, revealing that top-tier open-source models are closing the performance gap with proprietary models.

27Multimodal Large Language ModelsMLLMsHF ↗arXiv ↗
44

Paper2Web: Let's Make Your Paper Alive!

Yuhang Chen, Tianpeng Lv, Siyi Zhang +4 authors

Paper2Web is a benchmark and evaluation framework for academic webpage generation, featuring PWAgent, an autonomous pipeline that enhances content and layout through MCP tools, outperforming end-to-end baselines.

27Large Language Model (LLM)Paper2WebHF ↗arXiv ↗
46

Language Models Model Language

Łukasz Borchmann

The paper advocates for an empiricist approach to evaluating language models, emphasizing frequency of use over traditional theoretical frameworks.

26HF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号