---
title: Word Embeddings
description: Page on Vedang Vatsa's site: https://veda.ng/glossary/word-embeddings
canonical: https://veda.ng/glossary/word-embeddings
last_updated: 2026-10-03
type: text/markdown
---
# Word Embeddings

Source: https://veda.ng/glossary/word-embeddings
Author: Vedang Vatsa (https://veda.ng/about)

Word embeddings are numeric vectors for words that place meaning and relationships in a continuous space.

Each word maps to a list of real numbers, typically 50 to 300 dimensions. Words that appear in similar contexts sit close together. Vectors are learned from large text corpora and reflect statistical patterns, not hand-written rules. The classic check is the king-queen analogy: the vector from "king" to "queen" is roughly the same as from "man" to "woman." The algorithm finds that by reading millions of sentences.

Embeddings turn text into linear algebra. Classification, translation, sentiment analysis, and question answering run more efficiently on vectors than on raw strings. Search ranks by semantic similarity. Voice assistants parse intent. Recommenders match products to review text. GloVe, fastText, and BERT produce reusable embeddings for research and production.

A 50 to 300 dimensional vector per word is learned so co-occurring words have high similarity. Arithmetic such as king minus man plus woman landing near queen is a diagnostic, not a toy. GloVe counts global co-occurrence. fastText adds subword n-grams so rare and misspelled words still get vectors. BERT produces contextual embeddings so "bank" in a river and "bank" in finance differ. Search, assistants, sentiment, translation, QA, and recommenders consume those vectors instead of raw strings. Once you have linear algebra, cosine and nearest neighbors become the API. The 2013 word2vec paper showed that vector offsets can capture analogies such as king minus man plus woman.

Glossary index: https://veda.ng/glossary