TensorX
返回文献探索

Paper · arXiv 2410.24159

GPT or BERT: why not both?

Lucas Georges Gabriel Charpentier, David Samuel

14 upvotesOctober 31, 2024arXiv 预印本
AI 摘要

A hybrid model combining masked and causal language modeling within a transformer stack outperforms models using either paradigm alone.

masked language modelingcausal language modelingtransformer stackGPT-BERTBabyLM Challengepretraining process

Abstract

We present a simple way to merge masked language modeling with causal language modeling. This hybrid training objective results in a model that combines the strengths of both modeling paradigms within a single transformer stack: GPT-BERT can be transparently used like any standard causal or masked language model. We test the pretraining process that enables this flexible behavior on the BabyLM Challenge 2024. The results show that the hybrid pretraining outperforms masked-only or causal-only models. We openly release the models, training corpora and code.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
GPT or BERT: why not both? | TensorX