TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Apr 22 – Apr 28, 2024
本周最热262

Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Marah Abdin, Sam Ade Jacobs, Ammar Ahmad Awan +81 authors

Phi-3-mini, a compact 3.8 billion parameter language model, achieves competitive performance with larger models through an enhanced training dataset and alignment.

language modelMMLUMT-benchrobustnessHF ↗arXiv ↗

44 篇论文 · 按点赞排序

04

Multi-Head Mixture-of-Experts

Xun Wu, Shaohan Huang, Wenhui Wang +1 authors

MH-MoE enhances SMoE by splitting tokens into sub-tokens processed by diverse experts in parallel, improving expert activation, context understanding, and overfitting, and demonstrating effectiveness across various modeling tasks.

61Sparse Mixtures of Experts (SMoE)Multi-Head Mixture-of-Experts (MH-MoE)HF ↗arXiv ↗
06

Make Your LLM Fully Utilize the Context

Shengnan An, Zexiong Ma, Zeqi Lin +2 authors

FILM-7B enhances long-context processing by using information-intensive training with synthesized datasets, improving retrieval and performance on real-world tasks without compromising short-context performance.

54information-intensive traininglong-context trainingHF ↗arXiv ↗
10

FlowMind: Automatic Workflow Generation with LLMs

Zhen Zeng, William Watson, Nicole Cho +4 authors

FlowMind uses Large Language Models with a generic prompt recipe to generate automatic workflows, addressing spontaneous tasks and ensuring data integrity, and it is evaluated using a new financial dataset NCEN-QA.

34Large Language ModelsGenerative Pretrained TransformerHF ↗arXiv ↗
11

Pegasus-v1 Technical Report

Raehyuk Jung, Hyojun Go, Jaehyuk Yi +41 authors

Pegasus-1 is a multimodal language model designed for video content comprehension, handling spatiotemporal information and demonstrated in benchmarks for video conversation, zero-shot video question answering, and video summarization.

33multimodal language modelvideo content understandingHF ↗arXiv ↗
12

FlashSpeech: Efficient Zero-Shot Speech Synthesis

Zhen Ye, Zeqian Ju, Haohe Liu +10 authors

FlashSpeech, using latent consistency model and adversarial consistency training, achieves fast and high-quality zero-shot speech synthesis with enhanced prosody generation.

31latent consistency modeladversarial consistency trainingHF ↗arXiv ↗
14

TextSquare: Scaling up Text-Centric Visual Instruction Tuning

Jingqun Tang, Chunhui Lin, Zhen Zhao +13 authors

Text-centric VQA performance is significantly improved by a large, high-quality instruction-tuning dataset, Square-10M, generated through a four-step process, leading to state-of-the-art results on multiple benchmarks.

30Multimodal Large Language ModelsMLLMsHF ↗arXiv ↗
17

SnapKV: LLM Knows What You are Looking for Before Generation

Yuhong Li, Yingbing Huang, Bowen Yang +6 authors

SnapKV reduces KV cache size and computational overhead in LLMs by compressing the cache based on observed attention patterns, improving decoding speed and memory efficiency without significant performance loss.

26Large Language ModelsKey-Value cacheHF ↗arXiv ↗
21

A Multimodal Automated Interpretability Agent

Tamar Rott Shaham, Sarah Schwettmann, Franklin Wang +4 authors

MAIA, a multimodal automated interpretability agent, uses neural models to perform feature interpretation and failure mode discovery for other models, demonstrating comparable results to human experimenters and aiding in reducing sensitivity to spurious features and identifying potential misclassifications.

22neural modelsfeature interpretationHF ↗arXiv ↗
26

Tele-FLM Technical Report

Xiang Li, Yiqun Yao, Xin Jiang +17 authors

An open-sourced 52B multilingual large language model, Tele-FLM, features efficient pre-training and superior language capabilities, and is comparable to larger models in terms of performance metrics.

18LLMslarge language modelsHF ↗arXiv ↗
29

NeRF-XL: Scaling NeRFs with Multiple GPUs

Ruilong Li, Sanja Fidler, Angjoo Kanazawa +1 authors

NeRF-XL enables the scalable training and rendering of large-capacity NeRFs across multiple GPUs, demonstrating improved reconstruction quality and speed on extensive datasets.

14Neural Radiance Fields (NeRFs)multi-GPU approachesHF ↗arXiv ↗
30

Music Consistency Models

Zhengcong Fei, Mingyuan Fan, Junshi Huang

Music Consistency Models (MusicCM) efficiently synthesize high-quality music by leveraging consistency models, reducing computational requirements and enabling real-time application.

14consistency modelsMusic Consistency ModelsHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号