TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

September 2023

50 篇论文 · 按点赞排序

32

Qwen Technical Report

Jinze Bai, Shuai Bai, Yunfei Chu +45 authors

Qwen, a series of large language models, including chat and specialized coding and mathematics variants, exhibit superior performance across various tasks and outperform open-source models.

39LLMslarge language modelsHF ↗arXiv ↗
36

RMT: Retentive Networks Meet Vision Transformers

Qihang Fan, Huaibo Huang, Mingrui Chen +2 authors

The proposed RMT model, combining RetNet and Transformer architectures, introduces explicit spatial distance priors and coordinate-wise decomposition to achieve exceptional performance in computer vision tasks.

34TransformerRetentive Network (RetNet)HF ↗arXiv ↗
39

One Wide Feedforward is All You Need

Telmo Pessoa Pires, António V. Lopes, Yannick Assogba +1 authors

The study shows that the Feed Forward Network in Transformers is highly redundant and its reduction or sharing within the model can lead to improved accuracy and latency without significant loss of performance.

34Transformer architectureAttentionHF ↗arXiv ↗
40

Text-to-3D using Gaussian Splatting

Zilong Chen, Feng Wang, Huaping Liu

Gaussian Splatting based text-to-3D generation improves geometry accuracy and fidelity through progressive optimization, including geometry optimization and appearance refinement.

33Gaussian Splatting3D priorHF ↗arXiv ↗
43

Aligning Large Multimodal Models with Factually Augmented RLHF

Zhiqing Sun, Sheng Shen, Shengcao Cao +9 authors

The paper presents Factually Augmented RLHF to address multimodal misalignment in large multimodal models, significantly improving hallucination reduction and overall performance on vision-language tasks compared to existing methods.

32Reinforcement Learning from Human Feedback (RLHF)vision-language alignmentHF ↗arXiv ↗
44

Effective Long-Context Scaling of Foundation Models

Wenhan Xiong, Jingyu Liu, Igor Molybog +18 authors

Long-context LLMs achieve significant advancements in handling extended contexts through continual pretraining and efficient instruction tuning, surpassing existing models on various benchmarks.

31LLMslong-contextHF ↗arXiv ↗
45

SLiMe: Segment Like Me

Aliasghar Khani, Saeid Asgari Taghanaki, Aditya Sanghi +2 authors

SLiMe segments images at desired granularity using Stable Diffusion with minimal annotations and outperforms existing one-shot and few-shot segmentation methods.

31Stable Diffusionattention mapsHF ↗arXiv ↗
46

AudioSR: Versatile Audio Super-resolution at Scale

Haohe Liu, Ke Chen, Qiao Tian +2 authors

A diffusion-based generative model, AudioSR, achieves robust audio super-resolution across various audio types and bandwidths, enhancing generation quality for different audio models.

29diffusion-based generative modelaudio super-resolutionHF ↗arXiv ↗
47

Tracking Anything with Decoupled Video Segmentation

Ho Kei Cheng, Seoung Wug Oh, Brian Price +2 authors

DEVA, a decoupled approach using image-level segmentation and bi-directional temporal propagation, achieves favorable results in various data-scarce video segmentation tasks with reduced annotation and training costs.

29video segmentationtask-specificHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号