TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

403 篇论文 · 按点赞排序

272

Gaussian Head Avatar: Ultra High-fidelity Head Avatar via Dynamic Gaussians

Yuelang Xu, Benwang Chen, Zhe Li +4 authors

The method uses controllable 3D Gaussians and a fully learned MLP-based deformation field for high-fidelity 3D head avatar modeling under sparse views, employing geometry-guided initialization with implicit SDF and Deep Marching Tetrahedra for stability and high rendering quality.

273D GaussiansMLP-based deformation fieldHF ↗arXiv ↗
273

CapsFusion: Rethinking Image-Text Data at Scale

Qiying Yu, Quan Sun, Xiaosong Zhang +4 authors

CapsFusion is an advanced framework that improves multimodal pretraining data by combining web-based image-text pairs and synthetic captions, leading to enhanced model performance, sample efficiency, and scalability.

27multimodal modelszero-shotHF ↗arXiv ↗
274

FreeInit: Bridging Initialization Gap in Video Diffusion Models

Tianxing Wu, Chenyang Si, Yuming Jiang +2 authors

FreeInit addresses the temporal consistency and unnatural dynamics issues in diffusion-based video generation by refining spatial-temporal low-frequency components during inference, improving subject appearance and consistency.

27diffusion-based video generationtemporal consistencyHF ↗arXiv ↗
275

Zero-Shot Metric Depth with a Field-of-View Conditioned Diffusion Model

Saurabh Saxena, Junhwa Hur, Charles Herrmann +2 authors

A generic diffusion model with log-scale depth parameterization and FOV conditioning achieves state-of-the-art zero-shot metric depth estimation by handling indoor and outdoor scenes effectively and reducing relative error significantly.

27monocular depth estimationdiffusion modelHF ↗arXiv ↗
276

PolyLM: An Open Source Polyglot Large Language Model

Xiangpeng Wei, Haoran Wei, Huan Lin +15 authors

PolyLM, a multilingual LLM trained on 640 billion tokens, enhances multilingual capabilities through bilingual data and curriculum learning, outperforming other models on multilingual tasks while maintaining English performance.

27large language modelsmultilingual LLMHF ↗arXiv ↗
279

One-for-All: Generalized LoRA for Parameter-Efficient Fine-tuning

Arnav Chavan, Zhuang Liu, Deepak Gupta +2 authors

GLoRA, an advanced method for parameter-efficient fine-tuning, enhances LoRA with a generalized prompt module and modular layer-wise structure search, offering superior performance across diverse tasks with fewer parameters and computational costs.

26Generalized LoRAGLoRAHF ↗arXiv ↗
283

AgentBench: Evaluating LLMs as Agents

Xiao Liu, Hao Yu, Hanchen Zhang +19 authors

AgentBench is a multi-dimensional benchmark for evaluating LLMs as autonomous agents across various interactive environments, highlighting performance differences between commercial and open-source models.

26Large Language Models (LLMs)AgentBenchHF ↗arXiv ↗
285

MobileSAMv2: Faster Segment Anything to Everything

Chaoning Zhang, Dongshen Han, Sheng Zheng +3 authors

A new approach improves MobileSAM's efficiency for the SegEvery task by directly generating valid masks, boosting performance and reducing computation time.

26segment anythingsegmentation tasksHF ↗arXiv ↗
287

Dual-Stream Diffusion Net for Text-to-Video Generation

Binhui Liu, Xin Liu, Anbo Dai +3 authors

The dual-stream diffusion net (DSDN) enhances video consistency and smoothness in text-to-video generation by using separate content and motion diffusion streams with a cross-transformer interaction module and motion decomposer/combiner.

25dual-stream diffusion net (DSDN)diffusion streamsHF ↗arXiv ↗
289

Platypus: Quick, Cheap, and Powerful Refinement of LLMs

Ariel N. Lee, Cole J. Hunter, Nataniel Ruiz

A fine-tuned and merged family of large language models named Platypus, using LoRA modules and a curated dataset, achieves top performance on the Open LLM Leaderboard with reduced data and compute.

25Large Language ModelsLLMsHF ↗arXiv ↗
291

Controlling Text-to-Image Diffusion by Orthogonal Finetuning

Zeju Qiu, Weiyang Liu, Haiwen Feng +6 authors

Orthogonal Finetuning and Constrained Orthogonal Finetuning methods enhance text-to-image diffusion models by preserving hyperspherical energy and improving stability, leading to better generation quality and speed.

25diffusion modelsOrthogonal FinetuningHF ↗arXiv ↗
293

Contrastive Prefence Learning: Learning from Human Feedback without RL

Joey Hejna, Rafael Rafailov, Harshit Sikchi +4 authors

A new regret-based algorithm, Contrastive Preference Learning (CPL), learns optimal policies directly from human preferences without learning a reward function, addressing optimization challenges in Reinforcement Learning from Human Feedback (RLHF).

25Reinforcement Learning from Human FeedbackRLHFHF ↗arXiv ↗
294

How is ChatGPT's behavior changing over time?

Lingjiao Chen, Matei Zaharia, James Zou

The performance and behavior of GPT-3.5 and GPT-4 fluctuated significantly between March and June 2023 across various tasks, emphasizing the necessity for ongoing LLM quality monitoring.

25large language modelsLLMHF ↗arXiv ↗
295

Med-Flamingo: a Multimodal Medical Few-shot Learner

Michael Moor, Qian Huang, Shirley Wu +6 authors

Med-Flamingo, an adaptation of OpenFlamingo-9B for the medical domain, demonstrates few-shot capabilities in generative visual question answering with significant performance improvements as evaluated by clinicians.

24medical generative vision-language modelsfew-shot learnerHF ↗arXiv ↗
10 / 14

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号