TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Sep 4 – Sep 10, 2023
本周最热87

YaRN: Efficient Context Window Extension of Large Language Models

Bowen Peng, Jeffrey Quesnelle, Honglu Fan +1 authors

YaRN extends the context window of transformer-based language models like LLaMA with improved efficiency and performance.

Rotary Position EmbeddingsRoPEcompute-efficientcontext windowHF ↗arXiv ↗

42 篇论文 · 按点赞排序

02

Large Language Models as Optimizers

Chengrun Yang, Xuezhi Wang, Yifeng Lu +4 authors

OPRO, a method using large language models to optimize tasks described in natural language, outperforms human-designed prompts on various benchmark datasets.

79derivative-based algorithmsOptimization by PROmpting (OPRO)HF ↗arXiv ↗
06

One Wide Feedforward is All You Need

Telmo Pessoa Pires, António V. Lopes, Yannick Assogba +1 authors

The study shows that the Feed Forward Network in Transformers is highly redundant and its reduction or sharing within the model can lead to improved accuracy and latency without significant loss of performance.

34Transformer architectureAttentionHF ↗arXiv ↗
07

SLiMe: Segment Like Me

Aliasghar Khani, Saeid Asgari Taghanaki, Aditya Sanghi +2 authors

SLiMe segments images at desired granularity using Stable Diffusion with minimal annotations and outperforms existing one-shot and few-shot segmentation methods.

31Stable Diffusionattention mapsHF ↗arXiv ↗
08

Tracking Anything with Decoupled Video Segmentation

Ho Kei Cheng, Seoung Wug Oh, Brian Price +2 authors

DEVA, a decoupled approach using image-level segmentation and bi-directional temporal propagation, achieves favorable results in various data-scarce video segmentation tasks with reduced annotation and training costs.

29video segmentationtask-specificHF ↗arXiv ↗
14

CityDreamer: Compositional Generative Model of Unbounded 3D Cities

Haozhe Xie, Zhaoxi Chen, Fangzhou Hong +1 authors

CityDreamer is a compositional generative model that excels in generating realistic 3D cities by separating building generation from background objects using distinct modules and leveraging real-world datasets.

21compositional generative model3D city generationHF ↗arXiv ↗
15

FACET: Fairness in Computer Vision Evaluation Benchmark

Laura Gustafson, Chloe Rolland, Nikhila Ravi +5 authors

A benchmark named FACET is introduced to evaluate performance disparities of computer vision models across demographic attributes through extensive annotations and intersectional analysis.

20image classificationobject detectionHF ↗arXiv ↗
17

ImageBind-LLM: Multi-modality Instruction Tuning

Jiaming Han, Renrui Zhang, Wenqi Shao +14 authors

ImageBind-LLM uses a learnable bind network to enable large language models to follow multi-modal instructions through image-text alignment and a visual cache model.

17multi-modality instruction tuningImageBindHF ↗arXiv ↗
20

Efficient RLHF: Reducing the Memory Usage of PPO

Michael Santacroce, Yadong Lu, Han Yu +2 authors

Hydra-RLHF optimizes Reinforcement Learning with Human Feedback by integrating SFT and Reward models and dynamically disabling LoRA, reducing memory usage and latency while maintaining performance.

16Reinforcement Learning with Human FeedbackRLHFHF ↗arXiv ↗
21

Matcha-TTS: A fast TTS architecture with conditional flow matching

Shivam Mehta, Ruibo Tu, Jonas Beskow +2 authors

Matcha-TTS is a new encoder-decoder architecture for text-to-speech that uses optimal-transport conditional flow matching to produce high-quality outputs quickly and efficiently, outperforming existing models in speed, memory usage, and audio quality.

15Matcha-TTSencoder-decoder architectureHF ↗arXiv ↗
22

PromptTTS 2: Describing and Generating Voices with Text Prompt

Yichong Leng, Zhifang Guo, Kai Shen +12 authors

PromptTTS 2 addresses the challenges in text-to-speech systems by using a variation network to capture voice variability not expressed in text prompts and a prompt generation pipeline utilizing large language models to create high-quality text prompts.

15text-to-speechtext promptsHF ↗arXiv ↗
30

Gated recurrent neural networks discover attention

Nicolas Zucchet, Seijin Kobayashi, Yassir Akram +4 authors

Recent RNNs with linear recurrent layers and multiplicative gating can implement linear self-attention, similar to Transformers, as discovered through reverse-engineering of trained RNNs on in-context learning tasks.

10RNNsrecurrent neural networksHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号