TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

November 2023
本月最热253

GAIA: a benchmark for General AI Assistants

Grégoire Mialon, Clémentine Fourrier, Craig Swift +3 authors

GAIA benchmarks general AI assistants using real-world questions that challenge both reasoning and multi-modality handling, showcasing a significant gap between human and AI performance.

multi-modality handlingweb browsingtool-use proficiencyGPT-4HF ↗arXiv ↗

50 篇论文 · 按点赞排序

05

Orca 2: Teaching Small Language Models How to Reason

Arindam Mitra, Luciano Del Corro, Shweti Mahajan +12 authors

Orca 2 enhances smaller language models' reasoning abilities by teaching them diverse solution strategies, outperforming larger models on complex reasoning tasks.

78imitation learningreasoning techniquesHF ↗arXiv ↗
06

Make Pixels Dance: High-Dynamic Video Generation

Yan Zeng, Guoqiang Wei, Jiani Zheng +4 authors

PixelDance, a diffusion model-based approach, generates high-dynamic videos by incorporating image instructions for first and last frames alongside text instructions, surpassing current text-to-video methods in complexity and motion.

67diffusion modelsimage instructionsHF ↗arXiv ↗
11

LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents

Shilong Liu, Hao Cheng, Haotian Liu +10 authors

LLaVA-Plus, a general-purpose multimodal assistant, enhances large multimodal models by integrating pre-trained vision and vision-language models, performing tool-assisted tasks and improving interaction through direct image grounding.

52multimodal assistantpre-trained vision and vision-language modelsHF ↗arXiv ↗
12

LRM: Large Reconstruction Model for Single Image to 3D

Yicong Hong, Kai Zhang, Jiuxiang Gu +7 authors

A Large Reconstruction Model using a transformer-based architecture predicts 3D neural radiance fields from single images using massive multi-view training data.

52Large Reconstruction Modeltransformer-based architectureHF ↗arXiv ↗
13

Diffusion Model Alignment Using Direct Preference Optimization

Bram Wallace, Meihua Dang, Rafael Rafailov +7 authors

A method called Diffusion-DPO aligns text-to-image diffusion models to human preferences using direct optimization on comparison data, improving visual appeal and prompt alignment.

49Reinforcement Learning from Human Feedback (RLHF)human comparison dataHF ↗arXiv ↗
16

Drivable 3D Gaussian Avatars

Wojciech Zielonka, Timur Bagautdinov, Shunsuke Saito +3 authors

A new 3D controllable avatar model uses Gaussian splats for photorealistic rendering in real-time, employing cage deformations driven by joint angles and keypoints, outperforming existing methods.

47Gaussian splats3D Gaussian SplattingHF ↗arXiv ↗
17

Instant3D: Instant Text-to-3D Generation

Ming Li, Pan Zhou, Jia-Wei Liu +4 authors

A framework named Instant3D generates 3D objects from text prompts in under one second using a novel network and adaptive algorithms to enhance efficiency and quality.

47text-to-3D generationneural fieldHF ↗arXiv ↗
22

GPQA: A Graduate-Level Google-Proof Q&A Benchmark

David Rein, Betty Li Hou, Asa Cooper Stickland +5 authors

A dataset of extremely difficult multiple-choice questions challenges both experts and AI systems, facilitating the development of scalable oversight methods for AI-generated knowledge.

39GPQAmultiple-choice questionsHF ↗arXiv ↗
26

GLaMM: Pixel Grounding Large Multimodal Model

Hanoona Rasheed, Muhammad Maaz, Sahal Shaji +7 authors

GLaMM is a multimodal model that generates visually grounded language responses with object segmentation masks from both text and optional visual prompts.

36Large Multimodal ModelsLarge Language ModelsHF ↗arXiv ↗
28

Contrastive Chain-of-Thought Prompting

Yew Ken Chia, Guizhen Chen, Luu Anh Tuan +2 authors

Contrastive chain of thought, utilizing both valid and invalid reasoning examples, improves language model reasoning and generalization compared to conventional methods.

35chain of thoughtreasoningHF ↗arXiv ↗
29

ChatAnything: Facetime Chat with LLM-Enhanced Personas

Yilin Zhao, Xinbin Yuan, Shanghua Gao +4 authors

A framework for generating anthropomorphized personas with diverse voices and appearances from text descriptions using LLMs and generative models, with improved face landmark detection for automatic animation.

35LLM-based charactersin-context learningHF ↗arXiv ↗
30

OtterHD: A High-Resolution Multi-modality Model

Bo Li, Peiyuan Zhang, Jingkang Yang +3 authors

OtterHD-8B, an advanced multimodal model from Fuyu-8B, excels in processing high-resolution inputs and discerning detailed spatial relationships through MagnifierBench, an evaluation framework highlighting the importance of vision encoder flexibility.

34multimodal modelhigh-resolution visual inputsHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号