TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Dec 9 – Dec 15, 2024

50 篇论文 · 按点赞排序

33

JuStRank: Benchmarking LLM Judges for System Ranking

Ariel Gera, Odellia Boni, Yotam Perlitz +3 authors

A study evaluates LLM judges for system-level ranking in generative AI, assessing their quality, decisiveness, and bias through comparison with human-based rankings.

20LLM-based judgessystem-level rankingHF ↗arXiv ↗
35

Mobile Video Diffusion

Haitam Ben Yahia, Denis Korzhenkov, Ioannis Lelekas +2 authors

A mobile-optimized video diffusion model reduces computational demands and maintains quality through pruning and adversarial finetuning.

20spatio-temporal UNetMobileVDHF ↗arXiv ↗
36

OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations

Linke Ouyang, Yuan Qu, Hongbin Zhou +17 authors

Document content extraction is crucial in computer vision, especially for meeting the high-quality data needs of large language models (LLMs) and retrieval-augmented generation (RAG) technologies. However, current document parsing methods suffer from significant limitations in terms of diversity and comprehensive evaluation. To address these challenges, we introduce OmniDocBench, a novel multi-source benchmark designed to advance automated document content extraction. OmniDocBench includes a meticulously curated and annotated high-quality evaluation dataset comprising nine diverse document types, such as academic papers, textbooks, slides, among others. Our benchmark provides a flexible and comprehensive evaluation framework with 19 layout category labels and 14 attribute labels, enabling multi-level assessments across entire datasets, individual modules, or specific data types. Using OmniDocBench, we perform an exhaustive comparative analysis of existing modular pipelines and multimodal end-to-end methods, highlighting their limitations in handling document diversity and ensuring fair evaluation. OmniDocBench establishes a robust, diverse, and fair evaluation standard for the document content extraction field, offering crucial insights for future advancements and fostering the development of document parsing technologies. The codes and dataset is available in https://github.com/opendatalab/OmniDocBench.

20HF ↗arXiv ↗
38

Granite Guardian

Inkit Padhi, Manish Nagireddy, Giandomenico Cornacchia +19 authors

Granite Guardian models provide comprehensive risk detection for prompts and responses in large language models (LLM), addressing social bias, profanity, violence, sexual content, unethical behavior, jailbreaking, and retrieval-augmented generation (RAG) hallucination risks.

20Granite Guardian modelslarge language model (LLM)HF ↗arXiv ↗
42

StreamChat: Chatting with Streaming Video

Jihao Liu, Zhiding Yu, Shiyi Lan +5 authors

StreamChat enhances Large Multimodal Models' interaction with streaming video by updating visual context at each decoding step and using a crossattention-based architecture with a 3D-RoPE mechanism for temporal information encoding.

18Large Multimodal Modelsstreaming videoHF ↗arXiv ↗
43

MoViE: Mobile Diffusion for Video Editing

Adil Karjauv, Noor Fathima, Ioannis Lelekas +3 authors

Optimizations including a lightweight autoencoder, classifier-free guidance distillation, and an adversarial distillation scheme enable high-speed, high-quality video editing on mobile devices.

18diffusion-based video editinglightweight autoencoderHF ↗arXiv ↗
48

Gated Delta Networks: Improving Mamba2 with Delta Rule

Songlin Yang, Jan Kautz, Ali Hatamizadeh

Gated DeltaNet, combining gating and delta update mechanisms, outperforms existing models in various tasks and achieves better performance through hybrid architectures with additional attention layers.

17Linear Transformersstandard TransformersHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号