TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

本月最热205

3D Gaussian Splatting for Real-Time Radiance Field Rendering

Bernhard Kerbl, Georgios Kopanas, Thomas Leimkühler +1 authors

A method using 3D Gaussians for scene representation and optimized rendering allows high-quality, real-time novel-view synthesis at 1080p resolution.

radiance fieldnovel-view synthesis3D Gaussiansvolumetric radiance fieldsHF ↗arXiv ↗

50 篇论文 · 按点赞排序

03

Self-Alignment with Instruction Backtranslation

Xian Li, Ping Yu, Chunting Zhou +5 authors

A scalable instruction-following language model is built using auto-labelling and iterative self-augmentation and self-curation, outperforming other LLaMa-based models on Alpaca.

43instruction backtranslationlanguage modelHF ↗arXiv ↗
06

LP-MusicCaps: LLM-Based Pseudo Music Captioning

SeungHeon Doh, Keunwoo Choi, Jongpil Lee +1 authors

A large-scale pseudo music caption dataset generated using LLMs improves music captioning performance over supervised models in zero-shot and transfer-learning settings.

39large language modelsLLMsHF ↗arXiv ↗
10

Learning to Model the World with Language

Jessy Lin, Yuqing Du, Olivia Watkins +4 authors

Dynalang, a multimodal agent, learns to predict future text and image representations using language hints to improve task performance and enrich its understanding.

36multimodal world modelself-supervised learningHF ↗arXiv ↗
11

TeCH: Text-guided Reconstruction of Lifelike Clothed Humans

Yangyi Huang, Hongwei Yi, Yuliang Xiu +4 authors

TeCH reconstructs high-fidelity 3D human models from single images using descriptive text prompts, a personalized Text-to-Image diffusion model, and a hybrid DMTet representation, outperforming existing methods.

35descriptive text promptsgarment parsing modelHF ↗arXiv ↗
12

LLaSM: Large Language and Speech Model

Yu Shu, Siwei Dong, Guangyao Chen +5 authors

LLaSM, an end-to-end trained large multi-modal speech-language model, enhances human interaction with AI by following speech-and-language instructions using cross-modal conversational abilities.

34multi-modal large language modelsvision-language multi-modal modelsHF ↗arXiv ↗
15

OctoPack: Instruction Tuning Code Large Language Models

Niklas Muennighoff, Qian Liu, Armel Zebaze +7 authors

Instruction tuning using Git commits improves performance on natural language and coding tasks compared to other benchmarks, with models achieving state-of-the-art results on expanded HumanEvalPack.

33instruction tuningcodeHF ↗arXiv ↗
17

Shepherd: A Critic for Language Model Generation

Tianlu Wang, Ping Yu, Xiaoqing Ellen Tan +7 authors

Shepherd, a small language model tuned for critique, outperforms or ties with larger models like ChatGPT in refining language model outputs using a high-quality feedback dataset.

33language modelcritiqueHF ↗arXiv ↗
23

AlphaStar Unplugged: Large-Scale Offline Reinforcement Learning

Michaël Mathieu, Sherjil Ozair, Srivatsan Srinivasan +21 authors

A benchmark called AlphaStar Unplugged is established for offline reinforcement learning using a dataset from StarCraft II, featuring unprecedented challenges and achieving high win rates with offline data.

29offline RL algorithmsbehavior cloningHF ↗arXiv ↗
28

AgentBench: Evaluating LLMs as Agents

Xiao Liu, Hao Yu, Hanchen Zhang +19 authors

AgentBench is a multi-dimensional benchmark for evaluating LLMs as autonomous agents across various interactive environments, highlighting performance differences between commercial and open-source models.

26Large Language Models (LLMs)AgentBenchHF ↗arXiv ↗
29

Dual-Stream Diffusion Net for Text-to-Video Generation

Binhui Liu, Xin Liu, Anbo Dai +3 authors

The dual-stream diffusion net (DSDN) enhances video consistency and smoothness in text-to-video generation by using separate content and motion diffusion streams with a cross-transformer interaction module and motion decomposer/combiner.

25dual-stream diffusion net (DSDN)diffusion streamsHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号