AniDoc: Animation Creation Made Easier
Yihao Meng, Hao Ouyang, Hanlin Wang +6 authors
AniDoc uses video diffusion models to automate colorization and in-betweening in 2D animation, improving efficiency by leveraging correspondence matching.
Explore · 每周精选
发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。
598 篇论文 · 按点赞排序
Yihao Meng, Hao Ouyang, Hanlin Wang +6 authors
AniDoc uses video diffusion models to automate colorization and in-betweening in 2D animation, improving efficiency by leveraging correspondence matching.
Wenyi Hong, Weihan Wang, Ming Ding +22 authors
The CogVLM2 family of visual language models, including CogVLM2, CogVLM2-Video, and GLM-4V, advances image and video understanding with enhanced architectures and training methods.
Santosh V. Patapati
A BiLSTM-based multi-modal model combines audio, facial, and text data using a GPT-4 model to classify depression, achieving state-of-the-art performance in binary classification tasks.
Soham De, Samuel L. Smith, Anushan Fernando +14 authors
Hawk and Griffin, models incorporating gated linear recurrences, achieve high performance with lower computational resources and better hardware efficiency compared to existing architectures.
Yuang Peng, Yuxin Cui, Haomiao Tang +7 authors
dreambench++ uses advanced multimodal GPT models to create human-aligned evaluations for generative models in image generation, improving assessment accuracy and efficiency.
Shuai Tan, Biao Gong, Xiang Wang +6 authors
Animate-X, based on latent diffusion models, generates high-quality videos for various character types by enhancing motion representation through the Pose Indicator and a new benchmark, demonstrating superior performance.
Hyeongmin Lee, Jin-Young Kim, Kyungjune Baek +18 authors
The study introduces TWLV-I, a novel video foundation model that outperforms existing models in video comprehension by constructing robust representations for both motion and appearance.
Junyou Li, Qin Zhang, Yangbin Yu +2 authors
A sampling-and-voting method enhances large language models' performance by increasing the number of agents, with effectiveness tied to task difficulty.
Qixun Wang, Xu Bai, Haofan Wang +2 authors
InstantID is a diffusion model-based solution for personalized image synthesis that uses a single facial image and integrates with pre-trained models while maintaining high face fidelity and efficiency.
Wenxuan Zhang, Hou Pong Chan, Yiran Zhao +9 authors
SeaLLMs 3 is a large language model optimized for Southeast Asian languages, offering state-of-the-art performance in various tasks while addressing safety and cultural considerations.
Shivalika Singh, Freddie Vargus, Daniel Dsouza +30 authors
The initiative builds a human-curated instruction-following dataset spanning 65 languages and creates the largest multilingual collection of instruction-following instances through templating and translating existing datasets across 114 languages, contributing datasets and platforms for participatory research.
Hugo Laurençon, Léo Tronchon, Victor Sanh
A synthetic dataset of HTML code and screenshots is introduced to enhance the ability of vision-language models to convert screenshots into HTML code.
Ruibin Yuan, Hanfeng Lin, Yi Wang +32 authors
ChatMusician, an open-source LLM with intrinsic musical abilities, outperforms GPT-4 in music generation and understanding tasks using text-compatible music notation.
Daniel Goldstein, Fares Obeid, Eric Alcaide +2 authors
A hybrid Linear Attention/Transformer model, GoldFinch, efficiently generates a compressed KV-Cache for large context lengths with linear time and space complexity, outperforming previous models.
Wenqiang Sun, Shuo Chen, Fangfu Liu +4 authors
DimensionX leverages video diffusion with spatial-temporal decoupling and trajectory-aware mechanisms to generate high-quality 3D and 4D scenes from single images.
Yupeng Zhou, Daquan Zhou, Ming-Ming Cheng +2 authors
StoryDiffusion, combining consistent self-attention and semantic motion prediction, enables generation of coherent and stable images and videos from textual descriptions.
DeepSeek-AI, Xiao Bi, Deli Chen +84 authors
DeepSeek LLM, an open-source language model project, develops a large dataset and employs SFT and DPO to achieve performance surpassing LLaMA-2 70B and GPT-3.5 in various benchmarks and open-ended evaluations.
Jeffrey Li, Alex Fang, Georgios Smyrnis +56 authors
DataComp for Language Models (DCLM) provides a benchmark with a standardized corpus and recipes to improve language models through controlled dataset experiments, showing the importance of data curation for performance and compute efficiency.
Chaofan Tao, Qian Liu, Longxu Dou +5 authors
Investigating the impact of vocabulary size on the scaling of large language models reveals that larger vocabularies improve performance when considering compute budgets, and current models often use suboptimal vocabulary sizes.
Praveen K Kanithi, Clément Christophe, Marco AF Pimentel +7 authors
MEDIC framework evaluates Large Language Models across five clinical dimensions to guide model selection in healthcare applications, identifying performance trade-offs and ensuring practical implementation.
Guosheng Dong, Da Pan, Yiding Sun +17 authors
A data processing pipeline and open-sourced details for training a large language model achieve competitive performance with commercial models.
Rongyao Fang, Chengqi Duan, Kun Wang +7 authors
PUMA, a unified multimodal large language model, addresses varying granularity demands in visual content generation tasks through the integration of multi-granular visual features.
Changyue Liao, Mo Sun, Zihan Yang +4 authors
Fuyou enables efficient fine-tuning of large models on low-end GPUs by optimizing SSD-CPU communication and computation, achieving higher GPU utilization compared to ZeRO-Infinity.
Qihang Fan, Quanzeng You, Xiaotian Han +5 authors
ViTAR enhances Vision Transformers' scalability across resolutions through dynamic token integration and fuzzy positional encoding, improving accuracy and reducing computational costs.
Junying Chen, Chi Gui, Anningzhe Gao +4 authors
Chain-of-Diagnosis (CoD) enhances interpretability in LLM-based medical diagnostics by providing a transparent reasoning pathway and developing DiagnosisGPT, which diagnoses a wide range of diseases with high accuracy and controllable rigor.
Jihwan Kim, Junoh Kang, Jinyoung Choi +1 authors
FIFO-Diffusion generates long text-conditional videos through iterative denoising with latent partitioning and lookahead strategies to minimize the training-inference gap.
Sean McLeish, Arpit Bansal, Alex Stein +8 authors
Transformers achieve state-of-the-art performance on large arithmetic tasks and other reasoning tasks by addressing positional tracking with embeddings and integrating architectural modifications.
Byung-Kwan Lee, Chae Won Kim, Beomchan Park +1 authors
Meteor, an efficient large language and vision model, enhances understanding and answering capabilities by embedding multifaceted rationales using the Mamba architecture, leading to improved vision language performance without increasing model size or using additional vision encoders.
Weiyun Wang, Shuibo Zhang, Yiming Ren +13 authors
NN-NIAH is a benchmark that evaluates MLLMs' comprehension of long multimodal documents across retrieval, counting, and reasoning tasks.
Xiaoyi Dong, Pan Zhang, Yuhang Zang +20 authors
InternLM-XComposer2, a vision-language model, excels in free-form text-image creation and understanding using Partial LoRA, surpassing existing models and matching GPT-4V and Gemini Pro in multimodal benchmarks.
北京市昌平区探索星信息技术及软件开发工作室
京ICP备2026059466号