TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

May 27 – Jun 2, 2024

50 篇论文 · 按点赞排序

32

GFlow: Recovering 4D World from Monocular Video

Shizun Wang, Xingyi Yang, Qiuhong Shen +2 authors

GFlow reconstructs 4D dynamic scenes and相机 poses from a single monocular video using Gaussian splatting and 2D priors, enabling novel view rendering and object tracking.

15monocular video4D reconstructionHF ↗arXiv ↗
40

GECO: Generative Image-to-3D within a SECOnd

Chen Wang, Jiatao Gu, Xiaoxiao Long +2 authors

GECO, a two-stage generative model using score distillation, achieves efficient high-quality 3D generation from images by addressing view inconsistency.

12score distillationmulti-view generative modelHF ↗arXiv ↗
42

EM Distillation for One-step Diffusion Models

Sirui Xie, Zhisheng Xiao, Diederik P Kingma +6 authors

EM Distillation stabilizes the distillation of diffusion models into one-step generators, enhancing perceptual quality and outperforming existing methods in FID scores and text-to-image tasks.

12diffusion modelssamplingHF ↗arXiv ↗
49

DevEval: A Manually-Annotated Code Generation Benchmark Aligned with Real-World Code Repositories

Jia Li, Ge Li, Yunfei Zhao +15 authors

How to evaluate the coding abilities of Large Language Models (LLMs) remains an open question. We find that existing benchmarks are poorly aligned with real-world code repositories and are insufficient to evaluate the coding abilities of LLMs. To address the knowledge gap, we propose a new benchmark named DevEval, which has three advances. (1) DevEval aligns with real-world repositories in multiple dimensions, e.g., code distributions and dependency distributions. (2) DevEval is annotated by 13 developers and contains comprehensive annotations (e.g., requirements, original repositories, reference code, and reference dependencies). (3) DevEval comprises 1,874 testing samples from 117 repositories, covering 10 popular domains (e.g., Internet, Database). Based on DevEval, we propose repository-level code generation and evaluate 8 popular LLMs on DevEval (e.g., gpt-4, gpt-3.5, StarCoder 2, DeepSeek Coder, CodeLLaMa). Our experiments reveal these LLMs' coding abilities in real-world code repositories. For example, in our experiments, the highest Pass@1 of gpt-4-turbo is only 53.04%. We also analyze LLMs' failed cases and summarize their shortcomings. We hope DevEval can facilitate the development of LLMs in real code repositories. DevEval, prompts, and LLMs' predictions have been released.

10repository-level code generationHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号