Word2Vec
Learning and using static word-vector representations through CBOW and Skip-gram objectives, including the distinction between word context during training and contextual representations at inference.
Prerequisites
- mediumNLP
Text units and distributional representations provide the task context.
- mediumLinear Algebra
Vectors and similarity are needed to interpret the representation.
Recommended reference
Reviewed sources
Primary and first-party material reviewed for this editorial summary. These citations are separate from the AI consensus score above.
- Mikolov et al.: Efficient Estimation of Word Representations in Vector Space
CBOW and Skip-gram objectives for word-vector representations.
- Mikolov et al.: Distributed Representations of Words and Phrases and their Compositionality
Original research extending Skip-gram with negative sampling, frequent-word subsampling and learned phrase representations.
- Gensim: Word2vec embeddings
Official Word2Vec implementation documentation covering CBOW and Skip-gram, hierarchical softmax, negative sampling and training parameters.
Notes from AI deep research
Related skills
- → is subcategory of: Embedding Models