TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Dec 11 – Dec 17, 2023
本周最热57

LLM360: Towards Fully Transparent Open-Source LLMs

Zhengzhong Liu, Aurick Qiao, Willie Neiswanger +25 authors

LLM360 initiative promotes full transparency and reproducibility in LLM training by open-sourcing training code, data, model checkpoints, and intermediate results.

Large Language ModelsLLaMAFalconMistralHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

StemGen: A music generation model that listens

Julian D. Parker, Janne Spijkervet, Katerina Kosta +6 authors

A transformer-based, non-autoregressive model generates musically coherent audio by responding to context, achieving audio quality comparable to text-conditioned models.

48non-autoregressivetransformer-basedHF ↗arXiv ↗
05

SparQ Attention: Bandwidth-Efficient LLM Inference

Luka Ribar, Ivan Chelombiev, Luke Hudlass-Galley +3 authors

SparQ Attention reduces memory bandwidth requirements in LLM attention blocks, enhancing inference throughput without accuracy loss.

40SparQ Attentiongenerative large language models (LLMs)HF ↗arXiv ↗
07

CogAgent: A Visual Language Model for GUI Agents

Wenyi Hong, Weihan Wang, Qingsong Lv +8 authors

CogAgent, a visual language model with strong GUI understanding and navigation capabilities, outperforms LLM-based methods in both PC and Android GUI tasks using only screenshots.

32visual language modelGUI understandingHF ↗arXiv ↗
10

FreeInit: Bridging Initialization Gap in Video Diffusion Models

Tianxing Wu, Chenyang Si, Yuming Jiang +2 authors

FreeInit addresses the temporal consistency and unnatural dynamics issues in diffusion-based video generation by refining spatial-temporal low-frequency components during inference, improving subject appearance and consistency.

26diffusion-based video generationtemporal consistencyHF ↗arXiv ↗
12

Photorealistic Video Generation with Diffusion Models

Agrim Gupta, Lijun Yu, Kihyuk Sohn +6 authors

A transformer-based diffusion model using causal encoding and window attention generates high-resolution, photorealistic videos, achieving state-of-the-art performance without classifier-free guidance and includes a cascade for text-to-video generation.

24transformer-based approachdiffusion modelingHF ↗arXiv ↗
13

VideoLCM: Video Latent Consistency Model

Xiang Wang, Shiwei Zhang, Han Zhang +4 authors

VideoLCM framework leverages consistency models for efficient video synthesis with minimal sampling steps, achieving high fidelity and temporal consistency.

23consistency modelslatent video diffusion modelsHF ↗arXiv ↗
16

VILA: On Pre-training for Visual Language Models

Ji Lin, Hongxu Yin, Wei Ping +7 authors

Enhanced pre-training methods improve visual language models by balancing LLM freezing, interleaved data use, and text-only instruction re-blending, leading to superior performance across benchmarks.

21visual language modelslarge language modelsHF ↗arXiv ↗
23

Mosaic-SDF for 3D Generative Models

Lior Yariv, Omri Puny, Natalia Neverova +2 authors

A new Mosaic-SDF representation for 3D shapes enables efficient, parameter-efficient, and parallelizable generation of 3D shapes using generative flow models.

16diffusion modelsflow-based generative modelsHF ↗arXiv ↗
25

Context Tuning for Retrieval Augmented Generation

Raviteja Anantha, Tharun Bethi, Danil Vodianik +1 authors

Context Tuning enhances Retrieval Augmented Generation by improving context retrieval, tool retrieval, and LLM-based planner accuracy using smart context signals and fusion techniques.

16Retrieval Augmented Generation (RAG)semantic searchHF ↗arXiv ↗
26

Pixel Aligned Language Models

Jiarui Xu, Xingyi Zhou, Shen Yan +5 authors

The developed vision-language model can perform location-aware tasks by handling inputs and outputs as pixel coordinates or locations, achieving state-of-the-art performance on several benchmarks.

15vision-language modelslocation-conditioned captioningHF ↗arXiv ↗
27

Foundation Models in Robotics: Applications, Challenges, and the Future

Roya Firoozi, Johnathan Tucker, Stephen Tian +12 authors

Pretrained foundation models, trained on internet-scale data, offer improved generalization and can enhance various aspects of robot autonomy, including perception, decision-making, and control, despite challenges in robotics-specific data, safety, and real-time execution.

15pretrained foundation modelsinternet-scale dataHF ↗arXiv ↗
30

Alignment for Honesty

Yuqing Yang, Ethan Chern, Xipeng Qiu +2 authors

A paper proposes a framework and metrics to enhance the honesty of large language models by training them to accurately assess and reveal their knowledge limits.

13alignmenthonestyHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号