TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jul 15 – Jul 21, 2024
本周最热176

Qwen2 Technical Report

An Yang, Baosong Yang, Binyuan Hui +55 authors

The Qwen2 series, comprising 0.5 to 72 billion parameter models, surpasses prior open models across language understanding, generation, multilingualism, coding, math, and reasoning, with exceptional performance in benchmarks like MMLU, GPQA, HumanEval, GSM8K, BBH, MT-Bench, Arena-Hard, and LiveCodeBench.

Mixture-of-Expertslanguage modelsmultimodal modelsMMLUHF ↗arXiv ↗

50 篇论文 · 按点赞排序

04

Qwen2-Audio Technical Report

Yunfei Chu, Jin Xu, Qian Yang +9 authors

Qwen2-Audio, a large-scale audio-language model, enhances instruction-following and audio analysis through natural language prompts and DPO optimization.

64audio-language modelpre-training processHF ↗arXiv ↗
12

Toto: Time Series Optimized Transformer for Observability

Ben Cohen, Emaad Khwaja, Kan Wang +4 authors

Toto, a Time Series Optimized Transformer for Observability, achieves state-of-the-art performance in observability and general-purpose forecasting using a vast dataset of time series data.

33Time Series Optimized TransformerobservabilityHF ↗arXiv ↗
15

Scaling Diffusion Transformers to 16 Billion Parameters

Zhengcong Fei, Mingyuan Fan, Changqian Yu +2 authors

DiT-MoE, a sparse diffusion Transformer with shared expert routing and expert-level balance loss, achieves competitive performance with dense networks in image generation while reducing computational load during inference.

26diffusion Transformersparse versionHF ↗arXiv ↗
16

GRUtopia: Dream General Robots in a City at Scale

Hanqing Wang, Jiahe Chen, Wensi Huang +19 authors

Project GRUtopia introduces GRScenes, GRResidents, and GRBench to simulate a comprehensive 3D interactive society for embodied AI, focusing on diverse environments and advanced robotic tasks.

24Sim2RealEmbodied AIHF ↗arXiv ↗
17

Ref-AVS: Refer and Segment Objects in Audio-Visual Scenes

Yaoting Wang, Peiwen Sun, Dongzhan Zhou +3 authors

A new task, Reference Audio-Visual Segmentation (Ref-AVS), is introduced to segment objects using multimodal cues, and a method leveraging these cues outperforms existing approaches in experiments.

23Reference Audio-Visual SegmentationRef-AVSHF ↗arXiv ↗
20

MUSCLE: A Model Update Strategy for Compatible LLM Evolution

Jessica Echterhoff, Fartash Faghri, Raviteja Vemulapalli +4 authors

The work provides evaluation metrics and a training strategy to minimize inconsistencies and negative flips in Large Language Model updates, ensuring better compatibility with prior model versions for downstream tasks.

22Large Language Models (LLMs)generative tasksHF ↗arXiv ↗
21

Scaling Granite Code Models to 128K Context

Matt Stallone, Vaibhav Saxena, Leonid Karlinsky +19 authors

Granite code models with extended context windows up to 128K tokens are achieved through lightweight continual pretraining and finetuning with increased RoPE frequencies and length-upsampled data, showing improvements in long-context tasks without degrading performance on standard benchmarks.

21Granite code modelsRoPEHF ↗arXiv ↗
22

Shape of Motion: 4D Reconstruction from a Single Video

Qianqian Wang, Vickie Ye, Hang Gao +3 authors

A monocular dynamic reconstruction method using SE3 motion bases and data-driven priors achieves state-of-the-art performance in 3D/2D motion estimation and novel view synthesis.

20monocular dynamic reconstructionSE3 motion basesHF ↗arXiv ↗
24

H2O-Danube3 Technical Report

Pascal Pfeiffer, Philipp Singer, Yauhen Babakhin +3 authors

H2O-Danube3, a series of small language models, achieves high performance across various benchmarks and is efficient for local inference on smartphones.

19language modelspre-trainedHF ↗arXiv ↗
28

Patch-Level Training for Large Language Models

Chenze Shao, Fandong Meng, Jie Zhou

Patch-level training for LLMs reduces computational costs by compressing multiple tokens into patches, decreasing sequence length while maintaining model performance.

17Large Language Models (LLMs)token-level trainingHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号