TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Aug 12 – Aug 18, 2024
本周最热126

The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

Chris Lu, Cong Lu, Robert Tjarko Lange +3 authors

The AI Scientist is a framework that enables independent scientific research through automatic idea generation, experimentation, and paper writing using large language models.

diffusion modelingtransformer-based language modelinglearning dynamicsautomated reviewerHF ↗arXiv ↗

50 篇论文 · 按点赞排序

05

Imagen 3

Imagen-Team-Google, Jason Baldridge, Jakob Bauer +248 authors

Imagen 3, a latent diffusion model, generates high-quality images from text prompts and outperforms state-of-the-art models while addressing safety and representation issues.

62latent diffusion modelHF ↗arXiv ↗
07

Med42-v2: A Suite of Clinical LLMs

Clément Christophe, Praveen K Kanithi, Tathagata Raha +2 authors

Med42-v2 enhances Llama3 with clinical data to improve performance in healthcare settings, outperforming generic models across medical benchmarks.

52large language modelsLLMsHF ↗arXiv ↗
08

VITA: Towards Open-Source Interactive Omni Multimodal LLM

Chaoyou Fu, Haojia Lin, Zuwei Long +12 authors

VITA, an open-source Multimodal Large Language Model, excels in processing Video, Image, Text, and Audio with seamless interaction, showcasing advancements in multimodal understanding and human-computer interaction.

50Multimodal Large Language ModelMixtralHF ↗arXiv ↗
14

Layerwise Recurrent Router for Mixture-of-Experts

Zihan Qiu, Zeyu Huang, Shuang Cheng +4 authors

Layerwise Recurrent Router for Mixture-of-Experts (RMoE) improves MoE models by leveraging GRU-based cross-layer dependencies, enhancing expert selection and diversity without significant computational cost.

32Large language models (LLMs)Mixture-of-Experts (MoE)HF ↗arXiv ↗
17

Towards flexible perception with visual memory

Robert Geirhos, Priyank Jaini, Austin Stone +5 authors

A visual memory system that combines deep neural networks with databases to offer flexible data addition, removal, and interpretable decision-making, enhancing knowledge representation in visual models.

22deep neural networkspre-trained embeddingHF ↗arXiv ↗
19

Generative Photomontage

Sean J. Liu, Nupur Kumari, Ariel Shamir +1 authors

A framework using Generative Photomontage allows users to composite images from different parts of generated images, enhancing flexibility and quality in text-to-image models with new blending techniques.

20ControlNetbrush stroke interfaceHF ↗arXiv ↗
26

FuseChat: Knowledge Fusion of Chat Models

Fanqi Wan, Longguang Zhong, Ziyi Yang +2 authors

FuseChat integrates diverse chat LLMs using lightweight fine-tuning and parameter-space merging, resulting in a highly capable model competitive with larger LLMs.

15knowledge fusionchat LLMsHF ↗arXiv ↗
28

DeepSpeak Dataset v1.0

Sarah Barrington, Matyas Bohacek, Hany Farid

We describe a large-scale dataset--{\em DeepSpeak}--of real and deepfake footage of people talking and gesturing in front of their webcams. The real videos in this first version of the dataset consist of 9 hours of footage from 220 diverse individuals. Constituting more than 25 hours of footage, the fake videos consist of a range of different state-of-the-art face-swap and lip-sync deepfakes with natural and AI-generated voices. We expect to release future versions of this dataset with different and updated deepfake technologies. This dataset is made freely available for research and non-commercial uses; requests for commercial use will be considered.

14deepfakeface-swapHF ↗arXiv ↗
29

Aquila2 Technical Report

Bo-Wen Zhang, Liangdong Wang, Jijie Li +6 authors

The Aquila2 series of bilingual models, using the HeuriMentor framework, achieve competitive performance on English and Chinese benchmarks with efficient training and data management.

14HeuriMentorHM SystemHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号