TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

September 2023
本月最热106

MVDream: Multi-view Diffusion for 3D Generation

Yichun Shi, Peng Wang, Jianglong Ye +3 authors

MVDream generates geometrically consistent multi-view images from text prompts using pre-trained image diffusion models and Score Distillation Sampling, improving 3D generation stability and supporting personalized generation.

multi-view diffusion modeltext promptimage diffusion modelslarge-scale web datasetsHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

Textbooks Are All You Need II: phi-1.5 technical report

Yuanzhi Li, Sébastien Bubeck, Ronen Eldan +3 authors

A new 1.3 billion parameter Transformer-based language model, phi-1.5, demonstrates comparable performance to much larger models on common sense reasoning and complex tasks despite the absence of web data.

92Transformer-based language modelsTinyStoriesHF ↗arXiv ↗
05

Vision Transformers Need Registers

Timothée Darcet, Maxime Oquab, Julien Mairal +1 authors

Additional input tokens in Vision Transformers mitigate artifacts in feature maps, enhancing performance and enabling smoother visual processing.

86Transformersfeature mapsHF ↗arXiv ↗
07

Language Modeling Is Compression

Grégoire Delétang, Anian Ruoss, Paul-Ambroise Duquenne +9 authors

Large language models demonstrate strong compression capabilities, outperforming domain-specific compressors and offering new insights into scaling laws, tokenization, and in-context learning.

85self-supervised modelslarge language modelsHF ↗arXiv ↗
09

CodePlan: Repository-level Coding using LLMs and Planning

Ramakrishna Bairi, Atharv Sonwane, Aditya Kanade +6 authors

CodePlan automates repository-level coding tasks, such as package migration and temporal code edits, using a planning framework that leverages LLMs with context derived from code repositories and change analysis.

80Large Language ModelsLLMsHF ↗arXiv ↗
10

NExT-GPT: Any-to-Any Multimodal LLM

Shengqiong Wu, Hao Fei, Leigang Qu +2 authors

NExT-GPT, an any-to-any Multimodal Large Language Model, combines LLMs with multimodal adaptors and diffusion decoders to generate content across various modalities, enhanced by modality-switching instruction tuning and a curated dataset.

79Multimodal Large Language ModelsMM-LLMsHF ↗arXiv ↗
11

Large Language Models as Optimizers

Chengrun Yang, Xuezhi Wang, Yifeng Lu +4 authors

OPRO, a method using large language models to optimize tasks described in natural language, outperforms human-designed prompts on various benchmark datasets.

79derivative-based algorithmsOptimization by PROmpting (OPRO)HF ↗arXiv ↗
14

FreeU: Free Lunch in Diffusion U-Net

Chenyang Si, Ziqi Huang, Yuming Jiang +1 authors

A method called FreeU improves diffusion U-Net models' generation quality by re-weighting skip connections and backbone features without additional training.

66diffusion U-NetU-NetHF ↗arXiv ↗
15

DreamLLM: Synergistic Multimodal Comprehension and Creation

Runpei Dong, Chunrui Han, Yuang Peng +11 authors

DreamLLM, a framework for Multimodal Large Language Models, directly samples in the multimodal space to enhance comprehension and creation synergy, enabling free-form interleaved content generation.

60generative modelingmultimodal spaceHF ↗arXiv ↗
16

AnyMAL: An Efficient and Scalable Any-Modality Augmented Language Model

Seungwhan Moon, Andrea Madotto, Zhaojiang Lin +10 authors

AnyMAL is a unified model that processes multiple modalities (text, image, video, audio, IMU) and generates text, achieving top performance in multimodal tasks through a pre-trained aligner and fine-tuning with diverse instructions.

56Any-Modality Augmented Language ModelAnyMALHF ↗arXiv ↗
17

Kosmos-2.5: A Multimodal Literate Model

Tengchao Lv, Yupan Huang, Jingye Chen +11 authors

Kosmos-2.5, a unified multimodal model, generates spatially-aware and structured text from text-intensive images using a Transformer architecture and task-specific prompts.

56multimodal literate modelmachine readingHF ↗arXiv ↗
19

Large-Scale Automatic Audiobook Creation

Brendan Walsh, Mark Hamilton, Greg Newby +8 authors

A system leverages neural text-to-speech to automatically generate high-quality, customizable audiobooks from e-books, significantly expanding access to literature.

55neural text-to-speechHF ↗arXiv ↗
20

Generative Image Dynamics

Zhengqi Li, Richard Tucker, Noah Snavely +1 authors

A frequency-coordinated diffusion sampling process is used to predict long-term motion representations for still images, enabling dynamic video creation and interactive scene manipulation.

54frequency-coordinated diffusion sampling processneural stochastic motion textureHF ↗arXiv ↗
28

Agents: An Open-source Framework for Autonomous Language Agents

Wangchunshu Zhou, Yuchen Eleanor Jiang, Long Li +14 authors

Agents is an open-source library that facilitates the creation, customization, and deployment of autonomous language agents with features like planning, memory, and tool usage, aiming to make advances in large language models accessible to a broader audience.

43large language modelsautonomous language agentsHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号