BERT
Adapting bidirectional encoder-only Transformer models pretrained with masked-language objectives to language-understanding tasks.
Also searchable as: BERTs, Bidirectional Encoder Representations from Transformers
Prerequisites
- mediumTransformer Architecture
Attention, positional representations and encoder blocks are needed to understand BERT.
Recommended reference
Reviewed sources
Primary and first-party material reviewed for this editorial summary. These citations are separate from the AI consensus score above.
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Original paper on bidirectional Transformer language representations and adaptation to language-understanding tasks.
- Hugging Face Transformers — BERT
Official model implementation documentation covering masked-token pretraining, tokenization and downstream task heads.
- Google Research — BERT code and pretrained models
Original project repository with pretrained model use, fine-tuning workflows and the distinction between pretraining and downstream adaptation.
Notes from AI deep research
Related skills
- ← is subcategory of: RoBERTa
- → is subcategory of: Transformer Architecture