TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Dec 4 – Dec 10, 2023
本周最热152

Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Albert Gu, Tri Dao

Mamba, a novel SSM-based model, outperforms Transformers in inference speed and scalability across various modalities by selectively propagating information and using efficient hardware-aware algorithms.

Transformer architectureattention modulelinear attentiongated convolutionHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

Magicoder: Source Code Is All You Need

Yuxiang Wei, Zhe Wang, Jiawei Liu +2 authors

Magicoder, using OSS-Instruct to incorporate open-source code snippets, achieves superior performance on coding benchmarks while reducing bias in synthetic data generation.

83Large Language ModelsLLMsHF ↗arXiv ↗
04

Kandinsky 3.0 Technical Report

Vladimir Arkhipkin, Andrei Filatov, Viacheslav Vasilev +6 authors

Kandinsky 3.0, a large-scale text-to-image model based on latent diffusion, improves quality and realism through a larger architecture and advanced text understanding.

45latent diffusionU-NetHF ↗arXiv ↗
08

Alpha-CLIP: A CLIP Model Focusing on Wherever You Want

Zeyi Sun, Ye Fang, Tong Wu +6 authors

Alpha-CLIP enhances CLIP by adding an auxiliary alpha channel for attentive region suggestion, enabling precise control over image content across various tasks like open-world recognition and multimodal generation.

33CLIPContrastive Language-Image Pre-trainingHF ↗arXiv ↗
10

Analyzing and Improving the Training Dynamics of Diffusion Models

Tero Karras, Miika Aittala, Jaakko Lehtinen +3 authors

Modifications to network layers in the ADM diffusion model architecture improve training stability and synthesis quality, reducing FID from 2.41 to 1.81, and a novel method for post-hoc EMA parameter tuning is introduced.

32diffusion modelsADM diffusion modelHF ↗arXiv ↗
11

FaceStudio: Put Your Face Everywhere in Seconds

Yuxuan Yan, Chi Zhang, Rui Wang +3 authors

A hybrid guidance framework for identity-preserving image synthesis efficiently generates stylistic portraits by combining stylized images, facial images, and textual prompts.

32identity-preserving synthesisTextual InversionHF ↗arXiv ↗
12

Relightable Gaussian Codec Avatars

Shunsuke Saito, Gabriel Schwartz, Tomas Simon +2 authors

Relightable Gaussian Codec Avatars model high-fidelity head avatars with real-time relighting capabilities using 3D Gaussians for geometry and learnable radiance transfer for appearance.

313D Gaussiansrelightable appearance modelHF ↗arXiv ↗
14

Gaussian Head Avatar: Ultra High-fidelity Head Avatar via Dynamic Gaussians

Yuelang Xu, Benwang Chen, Zhe Li +4 authors

The method uses controllable 3D Gaussians and a fully learned MLP-based deformation field for high-fidelity 3D head avatar modeling under sparse views, employing geometry-guided initialization with implicit SDF and Deep Marching Tetrahedra for stability and high rendering quality.

263D GaussiansMLP-based deformation fieldHF ↗arXiv ↗
17

SeaLLMs -- Large Language Models for Southeast Asia

Xuan-Phi Nguyen, Wenxuan Zhang, Xin Li +9 authors

SeaLLMs, a series of Llama-2 based models tailored for Southeast Asian languages, demonstrate superior performance across linguistic tasks and outperform ChatGPT-3.5 in non-Latin languages while being lightweight and cost-effective.

24SeaLLMsLlama-2HF ↗arXiv ↗
18

Controllable Human-Object Interaction Synthesis

Jiaman Li, Alexander Clegg, Roozbeh Mottaghi +3 authors

CHOIS uses a conditional diffusion model to generate synchronized human and object motion in 3D scenes, guided by language descriptions and waypoints, with improvements for object geometry alignment and contact constraints.

23conditional diffusion modelobject motionHF ↗arXiv ↗
20

DeepCache: Accelerating Diffusion Models for Free

Xinyin Ma, Gongfan Fang, Xinchao Wang

DeepCache, a training-free method, accelerates diffusion models by caching and reusing features across denoising stages, achieving significant speedups with minimal quality degradation.

23diffusion modelssequential denoisingHF ↗arXiv ↗
26

Segment and Caption Anything

Xiaoke Huang, Jianfeng Wang, Yansong Tang +5 authors

A method enhances the Segment Anything Model (SAM) with regional caption generation using a lightweight query-based feature mixer, enabling efficient use of weak supervision pretraining on publicly available datasets.

21Segment Anything ModelSAMHF ↗arXiv ↗
28

Beyond Surface: Probing LLaMA Across Scales and Layers

Nuo Chen, Ning Wu, Shining Liang +4 authors

An analysis of LLaMA focusing on reasoning and computation through multiple-choice tasks reveals layer-specific knowledge and computational capabilities without significant enhancement from larger model sizes.

19Large Language Models (LLMs)LLaMAHF ↗arXiv ↗
29

AnimateZero: Video Diffusion Models are Zero-Shot Image Animators

Jiwen Yu, Xiaodong Cun, Chenyang Qi +4 authors

AnimateZero enhances pre-trained text-to-video diffusion models with precise appearance and motion control capabilities, enabling finer control over the generation process and supporting new applications like interactive video generation.

18text-to-video diffusion modelsappearance controlHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号