TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Feb 12 – Feb 18, 2024

50 篇论文 · 按点赞排序

34

Rolling Diffusion Models

David Ruhe, Jonathan Heek, Tim Salimans +1 authors

Rolling Diffusion enhances temporal data generation by progressively corrupting frames with increasing noise, outperforming standard diffusion in complex tasks like video prediction and fluid dynamics forecasting.

13diffusion modelsRolling DiffusionHF ↗arXiv ↗
35

Computing Power and the Governance of Artificial Intelligence

Girish Sastry, Lennart Heim, Haydn Belfield +16 authors

Computing power, or "compute," is crucial for the development and deployment of artificial intelligence (AI) capabilities. As a result, governments and companies have started to leverage compute as a means to govern AI. For example, governments are investing in domestic compute capacity, controlling the flow of compute to competing countries, and subsidizing compute access to certain sectors. However, these efforts only scratch the surface of how compute can be used to govern AI development and deployment. Relative to other key inputs to AI (data and algorithms), AI-relevant compute is a particularly effective point of intervention: it is detectable, excludable, and quantifiable, and is produced via an extremely concentrated supply chain. These characteristics, alongside the singular importance of compute for cutting-edge AI models, suggest that governing compute can contribute to achieving common policy objectives, such as ensuring the safety and beneficial use of AI. More precisely, policymakers could use compute to facilitate regulatory visibility of AI, allocate resources to promote beneficial outcomes, and enforce restrictions against irresponsible or malicious AI development and usage. However, while compute-based policies and technologies have the potential to assist in these areas, there is significant variation in their readiness for implementation. Some ideas are currently being piloted, while others are hindered by the need for fundamental research. Furthermore, naive or poorly scoped approaches to compute governance carry significant risks in areas like privacy, economic impacts, and centralization of power. We end by suggesting guardrails to minimize these risks from compute governance.

13HF ↗arXiv ↗
41

Scaling Laws for Fine-Grained Mixture of Experts

Jakub Krajewski, Jan Ludziejewski, Kamil Adamczewski +9 authors

MoE models reduce computational cost through fine-grained scaling controlled by a hyperparameter called granularity, outperforming dense Transformers across varying scales and budgets.

13Mixture of Experts (MoE)scaling propertiesHF ↗arXiv ↗
42

Hierarchical State Space Models for Continuous Sequence-to-Sequence Modeling

Raunaq Bhirangi, Chenyu Wang, Venkatesh Pattabiraman +4 authors

Hierarchical State-Space Models (HiSS) stack structured state-space models to create a temporal hierarchy, outperforming sequence models like causal Transformers, LSTMs, S4, and Mamba in continuous sequential prediction across real-world sensor datasets.

12Hierarchical State-Space ModelsHiSSHF ↗arXiv ↗
44

Learning Continuous 3D Words for Text-to-Image Generation

Ta-Ying Cheng, Matheus Gadelha, Thibault Groueix +4 authors

Continuous 3D Words enable fine-grained control of 3D attributes in text-to-image models, allowing users to adjust image generation through continuous sliders alongside text prompts.

11Continuous 3D Wordstext-to-image modelsHF ↗arXiv ↗
45

LiRank: Industrial Large Scale Ranking Models at LinkedIn

Fedor Borisyuk, Mingzhou Zhou, Qingquan Song +31 authors

LiRank, a large-scale ranking framework at LinkedIn, integrates advanced modeling techniques and optimization methods to enhance performance in feed ranking, job recommendations, and ads CTR prediction.

11Residual DCNattentionHF ↗arXiv ↗
46

Model Editing with Canonical Examples

John Hewitt, Sarah Chen, Lanruo Lora Xie +3 authors

Model editing using canonical examples improves performance with LoRA and sense finetuning on Pythia and Backpack language models.

11LoRAfull finetuningHF ↗arXiv ↗
50

SubGen: Token Generation in Sublinear Time and Memory

Amir Zandieh, Insu Han, Vahab Mirrokni +1 authors

A novel KV cache compression technique, SubGen, improves memory efficiency and performance in long-context token generation for large language models by using_online clustering and sampling.

10large language modelsmemory requirementsHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号