TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Dec 18 – Dec 24, 2023
本周最热265

LLM in a flash: Efficient Large Language Model Inference with Limited Memory

Keivan Alizadeh, Iman Mirzadeh, Dmitry Belenko +5 authors

Efficient inference for large language models on devices with limited DRAM by optimizing data transfer and access from flash memory.

large language models (LLMs)flash memoryDRAMinference cost modelHF ↗arXiv ↗

50 篇论文 · 按点赞排序

03

AppAgent: Multimodal Agents as Smartphone Users

Chi Zhang, Zhao Yang, Jiaxuan Liu +5 authors

A novel LLM-based multimodal agent learns to operate smartphone apps through autonomous exploration or imitation, demonstrating proficiency across diverse tasks.

54large language modelsmultimodal agentHF ↗arXiv ↗
05

Gemini: A Family of Highly Capable Multimodal Models

Gemini Team, Rohan Anil, Sebastian Borgeaud +939 authors

Gemini, a family of multimodal models, achieves state-of-the-art performance across various benchmarks, including human-expert performance on MMLU, through advanced cross-modal reasoning and language understanding.

50multimodal modelscross-modal reasoningHF ↗arXiv ↗
07

PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU

Yixin Song, Zeyu Mi, Haotong Xie +1 authors

PowerInfer, a high-speed LLM inference engine for personal computers, enhances efficiency using hotspot neuron analysis, GPU-CPU hybrid computation, adaptive predictors, and neuron-aware sparse operators, achieving performance close to server-grade GPUs.

46Large Language Model (LLM)inference engineHF ↗arXiv ↗
10

Generative Multimodal Models are In-Context Learners

Quan Sun, Yufeng Cui, Xiaosong Zhang +8 authors

A large-scale generative multimodal model with 37 billion parameters demonstrates strong few-shot in-context learning and achieves state-of-the-art performance on multimodal tasks through scaling-up and instruction tuning.

36task-agnostic in-context learninggenerative multimodal modelHF ↗arXiv ↗
13

DreamTuner: Single Image is Enough for Subject-Driven Generation

Miao Hua, Jiawei Liu, Fei Ding +3 authors

DreamTurner uses a novel approach by injecting reference information through subject encoders and self-subject-attention layers to enhance subject-driven image generation, balancing subject learning and model capabilities.

27diffusion-based modelstext-to-image generationHF ↗arXiv ↗
14

Zero-Shot Metric Depth with a Field-of-View Conditioned Diffusion Model

Saurabh Saxena, Junhwa Hur, Charles Herrmann +2 authors

A generic diffusion model with log-scale depth parameterization and FOV conditioning achieves state-of-the-art zero-shot metric depth estimation by handling indoor and outdoor scenes effectively and reducing relative error significantly.

27monocular depth estimationdiffusion modelHF ↗arXiv ↗
18

Paint3D: Paint Anything 3D with Lighting-Less Texture Diffusion Models

Xianfang Zeng, Xin Chen, Zhongqi Qi +6 authors

Paint3D uses a coarse-to-fine generative framework with specialized diffusion models to produce high-resolution, lighting-less UV texture maps for 3D meshes, addressing challenges of incomplete areas and illumination artifacts.

23coarse-to-fine generative frameworkdiffusion modelHF ↗arXiv ↗
19

VecFusion: Vector Font Generation with Diffusion

Vikas Thamizharasan, Difan Liu, Shantanu Agarwal +5 authors

A cascaded diffusion model combines raster and vector diffusion models, using a transformer architecture and novel vector representation to generate high-quality, complex vector fonts.

21neural architectureVecFusionHF ↗arXiv ↗
20

Point Transformer V3: Simpler, Faster, Stronger

Xiaoyang Wu, Li Jiang, Peng-Shuai Wang +6 authors

Point Transformer V3 improves efficiency and scaling in point cloud processing, achieving state-of-the-art performance by simplifying neighbor search and expanding receptive field.

21Point Transformer V3PTv3HF ↗arXiv ↗
21

Time is Encoded in the Weights of Finetuned Language Models

Kai Nylund, Suchin Gururangan, Noah A. Smith

Time vectors, derived by fine-tuning language models on specific time periods, enhance performance on text from those periods and can be interpolated to improve performance on future periods without further training.

20time vectorsfinetuningHF ↗arXiv ↗
24

GAvatar: Animatable 3D Gaussian Avatars with Implicit Mesh Learning

Ye Yuan, Xueting Li, Yangyi Huang +4 authors

A method combining Gaussian splatting with neural implicit fields and SDF-based mesh learning for generating realistic and animatable avatars from text prompts, achieving high-quality appearance and geometry with fast rendering.

19Gaussian splattingexplicit representationsHF ↗arXiv ↗
26

Rich Human Feedback for Text-to-Image Generation

Youwei Liang, Junfeng He, Gang Li +15 authors

Enriched human feedback is used to improve text-to-image generation through multimodal transformer training, enhancing quality and generalizability across different models.

19Text-to-ImageT2I generationHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号