TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Apr 1 – Apr 7, 2024

50 篇论文 · 按点赞排序

32

Measuring Style Similarity in Diffusion Models

Gowthami Somepalli, Anubhav Gupta, Kamal Gupta +5 authors

A new framework extracts style descriptors from images to attribute style in generated images to training data, with applications demonstrated in style retrieval and analysis for the Stable Diffusion model.

15generative modelstext-to-image modelsHF ↗arXiv ↗
34

PointInfinity: Resolution-Invariant Point Diffusion Models

Zixuan Huang, Justin Johnson, Shoubhik Debnath +2 authors

PointInfinity, a transformer-based point cloud diffusion model, efficiently generates high-resolution point clouds with state-of-the-art quality by using a fixed-size, resolution-invariant latent representation.

14point cloud diffusion modelstransformer-based architectureHF ↗arXiv ↗
36

Poro 34B and the Blessing of Multilinguality

Risto Luukkonen, Jonathan Burdge, Elaine Zosa +5 authors

A multilingual training approach on a large language model improves capabilities for small languages, translation, and generation in multiple languages.

13large language modelspretrainingHF ↗arXiv ↗
38

Streaming Dense Video Captioning

Xingyi Zhou, Anurag Arnab, Shyamal Buch +5 authors

A streaming dense video captioning model with a novel memory module and decoding algorithm improves state-of-the-art performance on multiple benchmarks while handling long videos efficiently.

12memory modulestreaming decoding algorithmHF ↗arXiv ↗
39

Condition-Aware Neural Network for Controlled Image Generation

Han Cai, Muyang Li, Zhuoyang Zhang +3 authors

A Condition-Aware Neural Network (CAN) method for image generation dynamically adjusts neural network weights based on input conditions, demonstrating significant improvements for diffusion transformer models and surpassing existing models in terms of FID score and computational efficiency.

12Condition-Aware Neural Network (CAN)condition-aware weight generation moduleHF ↗arXiv ↗
43

DiJiang: Efficient Large Language Models through Compact Kernelization

Hanting Chen, Zhicheng Liu, Xutao Wang +2 authors

DiJiang, a Frequency Domain Kernelization approach, reduces the computational load of pre-trained Transformers with little additional training, achieving comparable performance to vanilla Transformers with significantly lower costs and faster inference.

11linear attentionQuasi-Monte Carlo methodHF ↗arXiv ↗
46

Noise-Aware Training of Layout-Aware Language Models

Ritesh Sarkhel, Xiaoqi Ren, Lauro Beltrao Costa +5 authors

Noise-Aware Training (NAT) method improves efficiency and robustness in training named entity extractors using weakly labeled documents, outperforming transfer learning in terms of macro-F1 score and reducing human labeling effort.

9Noise-Aware TrainingNATHF ↗arXiv ↗
50

ST-LLM: Large Language Models Are Effective Temporal Learners

Ruyang Liu, Chen Li, Haoran Tang +3 authors

ST-LLM leverages large language models to effectively encode and understand videos for dialogue systems by integrating spatial-temporal modeling with dynamic masking and global-local input strategies, achieving state-of-the-art performance on VideoChatGPT-Bench and MVBench.

7Spatial-Temporal sequence modelingdynamic masking strategyHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号