AI Workspace Loading

We’re preparing your intelligent learning experience. Our AI systems are processing content, optimizing resources, and setting everything up for you.

Preparing Learning Paths...
AI Processing
Smart Automation
Learning Engine
Good things take a moment.

LearnLess.ai

LEARN LESS. UNDERSTAND MORE.
Intermediate4 min read

Semantic Similarity

Semantic similarity measures how close two pieces of content are in meaning, typically by comparing their embeddings.

Prerequisites

What Is Semantic Similarity?

Two pieces of text can mean nearly the same thing while sharing almost no exact words — "car won't start" and "vehicle fails to turn on" are semantically similar despite having no words in common. Semantic similarity is a measure of how close two pieces of content are in meaning, and embeddings are what make that measurable: text with similar meaning gets mapped to vectors that are close together in the embedding space.

Key Idea

Semantic similarity is measured between embeddings, not raw text — the embedding model is what turns "similar in meaning" into "close together as numbers."

Cosine Similarity: The Common Way to Measure It

The most common way to compare two embeddings is cosine similarity — it measures the angle between two vectors rather than their raw distance. Two vectors pointing in almost the same direction get a score near 1 (highly similar); vectors pointing in unrelated directions get a score near 0; opposite directions approach -1.

Text A

First Piece

The first piece of content being compared.

encoded as

Embedding A

Made Measurable

Turns meaning into comparable numbers.

compared via

Cosine Similarity

Angle, Not Distance

Measures direction between the two vectors.

against

Text B

Second Piece

May share no exact words with Text A.

encoded as

Embedding B

Also Made Measurable

Compared against Embedding A by angle.

Tip

Cosine similarity focuses on direction, not magnitude — two vectors can point the same way but have different lengths and still score as highly similar.

Code Example

cosine_similarity.py
import numpy as np

def cosine_similarity(a, b):
    a, b = np.array(a), np.array(b)
    return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))

query_embedding = embed("car won't start")
doc_embedding = embed("vehicle fails to turn on")

score = cosine_similarity(query_embedding, doc_embedding)
# score is close to 1.0 — high semantic similarity, despite no shared words

A Real-World Example

A support search tool built on keyword matching would miss a document about "vehicle fails to turn on" when a user searches "car won't start" — no shared words means no match. A search tool built on semantic similarity compares embeddings instead, and correctly finds the relevant document because the two phrases land close together in embedding space, even though the exact words are completely different.

Common Mistakes

  • Assuming semantic similarity replaces exact matching entirely

    Exact identifiers, codes, or rare terms are often matched better by keyword search — this is why hybrid search combines both.

  • Comparing embeddings from two different models

    Embeddings from different models generally live in different, incompatible vector spaces — comparisons only make sense within the same model.

  • Assuming a high similarity score means the content is fully relevant

    Two pieces of text can be topically similar without one actually answering the other — similarity is a signal, not a guarantee of relevance.

  • Ignoring which similarity metric a vector database uses

    Some systems default to a different distance metric (like dot product or Euclidean distance) — mismatching the metric your embeddings were designed for can degrade results.

Interview Question

What is semantic similarity, and how is cosine similarity used to measure it?

Semantic similarity measures how close two pieces of content are in meaning, which matters because two phrases can mean nearly the same thing while sharing no exact words. It's measured by comparing their embeddings rather than the raw text — content with similar meaning gets mapped to vectors that land close together. Cosine similarity is the most common way to compare two embeddings: it measures the angle between the vectors, giving a score near 1 for highly similar direction and near 0 for unrelated content, regardless of the vectors' magnitude.

What an interviewer may ask next

  • Why can two pieces of text be semantically similar despite sharing no exact words?
  • Why does cosine similarity focus on the angle between vectors rather than their length?
  • Why can't you meaningfully compare embeddings produced by two different embedding models?

Explain It in 30 Seconds

Semantic similarity measures how close two pieces of content are in meaning, by comparing their embeddings rather than their exact words. Cosine similarity is the most common way to do that comparison — it measures the angle between two embedding vectors, giving a score near 1 for very similar meaning and near 0 for unrelated content. This is what lets a search system match "car won't start" to "vehicle fails to turn on" even though they share no words.

On this page