primer.common.embedder

ConceptEmbedder: a deterministic, dependency-free stand-in for a real sentence-embedding model.

How it works (and why it's a fair teaching model):

  1. Tokenize the text (primer.common.text.tokenize).
  2. For each token, look it up in a tiny synonym lexicon (CONCEPTS). If the word belongs to a concept (e.g. "car" and "automobile" both belong to vehicle), add that concept's fixed random vector. Also add a smaller word-specific vector so synonyms are close but not identical.
  3. Words outside the lexicon (including IDs like "err-4012") get only their own random vector. They match themselves and nothing else.
  4. Mean-pool (sum) the token vectors and L2-normalize, exactly like a real bi-encoder does with mean pooling.

That reproduces the two behaviors that matter most in practice:

  • Dense retrieval finds meaning. "automobile reimbursement" is close to "car expense" even with zero shared words.
  • Dense retrieval blurs exact tokens. An error code is just one of many pooled token vectors, so it's a weak signal next to the concept words. That's why hybrid search (BM25 + dense) wins on enterprise data. See primer.ml.embeddings.retrieval.

Random high-dimensional vectors are nearly orthogonal (their cosine is about 0 plus or minus 1/sqrt(dim)), which is what makes "unrelated" words unrelated here.

Further reading:

on GitHub
  1"""
  2`ConceptEmbedder`: a deterministic, dependency-free stand-in for a real
  3sentence-embedding model.
  4
  5How it works (and why it's a fair teaching model):
  6
  71. Tokenize the text (`primer.common.text.tokenize`).
  82. For each token, look it up in a tiny synonym lexicon (`CONCEPTS`). If the
  9   word belongs to a concept (e.g. "car" and "automobile" both belong to
 10   `vehicle`), add that concept's fixed random vector. Also add a smaller
 11   word-specific vector so synonyms are close but not identical.
 123. Words outside the lexicon (including IDs like "err-4012") get only their
 13   own random vector. They match themselves and nothing else.
 144. Mean-pool (sum) the token vectors and L2-normalize, exactly like a real
 15   bi-encoder does with mean pooling.
 16
 17That reproduces the two behaviors that matter most in practice:
 18
 19* **Dense retrieval finds meaning.** "automobile reimbursement" is close to
 20  "car expense" even with zero shared words.
 21* **Dense retrieval blurs exact tokens.** An error code is just one of many
 22  pooled token vectors, so it's a weak signal next to the concept words.
 23  That's why hybrid search (BM25 + dense) wins on enterprise data. See
 24  `primer.ml.embeddings.retrieval`.
 25
 26Random high-dimensional vectors are nearly orthogonal (their cosine is
 27about 0 plus or minus 1/sqrt(dim)), which is what makes "unrelated" words
 28unrelated here.
 29
 30Further reading:
 31- Sentence-BERT (mean pooling, bi-encoders): https://arxiv.org/abs/1908.10084
 32- sentence-transformers docs: https://www.sbert.net/
 33- Random projections / near-orthogonality: https://en.wikipedia.org/wiki/Johnson%E2%80%93Lindenstrauss_lemma
 34"""
 35
 36from __future__ import annotations
 37
 38import hashlib
 39from functools import lru_cache
 40from typing import Iterable, Sequence
 41
 42import numpy as np
 43
 44from primer.common.text import tokenize
 45
 46# concept name -> words that express it. Keep it small and obvious.
 47CONCEPTS: dict[str, list[str]] = {
 48    "credential": ["password", "passwords", "passcode", "credentials", "login", "sign-in", "signin", "logon"],
 49    "reset": ["reset", "recover", "recovery", "forgot", "forgotten", "unlock", "locked", "lockout"],
 50    "mfa": ["2fa", "mfa", "authenticator", "two-factor", "otp", "verification"],
 51    "remote_access": ["vpn", "remote", "tunnel", "anyconnect"],
 52    "vehicle": ["car", "automobile", "vehicle", "auto", "mileage", "driving"],
 53    "money_back": ["reimburse", "reimbursement", "reimbursed", "expense", "expenses", "refund", "receipt", "receipts"],
 54    "time_off": ["pto", "vacation", "leave", "holiday", "holidays", "time-off", "sick"],
 55    "computer": ["laptop", "computer", "notebook", "macbook", "device", "hardware", "monitor"],
 56    "invoice": ["invoice", "invoices", "bill", "billing", "payment", "payments", "vendor", "vendors"],
 57    "travel": ["travel", "trip", "flight", "flights", "hotel", "airfare", "per-diem"],
 58    "error": ["error", "errors", "failure", "fail", "failed", "fails", "crash", "broken", "issue"],
 59    "email": ["email", "emails", "mail", "inbox", "outlook"],
 60    "threat": ["phishing", "suspicious", "scam", "malware", "attack"],
 61    "policy": ["policy", "policies", "rule", "rules", "requirement", "requirements", "guideline", "guidelines"],
 62    "new_hire": ["onboarding", "onboard", "hire", "hires", "orientation"],
 63    "reconcile": ["reconcile", "reconciliation", "match", "matching", "mismatch", "mismatches"],
 64    "salary": ["salary", "pay", "payroll", "paycheck", "compensation", "bonus"],
 65    "printer": ["printer", "printers", "print", "printing", "toner"],
 66}
 67
 68# Reverse index: word -> concept.
 69WORD_TO_CONCEPT: dict[str, str] = {w: c for c, ws in CONCEPTS.items() for w in ws}
 70
 71
 72def _seed(key: str) -> int:
 73    # Python's built-in hash() is salted per process; md5 is stable across runs.
 74    return int(hashlib.md5(key.encode()).hexdigest()[:16], 16)
 75
 76
 77@lru_cache(maxsize=None)
 78def _unit_vector(key: str, dim: int) -> np.ndarray:
 79    """A fixed pseudo-random unit vector for `key`. Same key -> same vector."""
 80    v = np.random.default_rng(_seed(key)).standard_normal(dim)
 81    v /= np.linalg.norm(v)
 82    v.setflags(write=False)  # cached arrays are shared; make them read-only
 83    return v
 84
 85
 86class ConceptEmbedder:
 87    """Deterministic toy bi-encoder. `encode(texts)` returns L2-normalized rows.
 88
 89    Args:
 90        dim: output dimensionality (real models use 384 to 3072).
 91        word_weight: how much each word's own identity counts relative to its
 92            concept. Lower means synonyms are closer together.
 93        name: a "model version" string. Two embedders with different names
 94            produce incompatible vector spaces, which lets
 95            `primer.ml.embeddings.operations` demonstrate why switching
 96            models means re-embedding everything.
 97    """
 98
 99    def __init__(self, dim: int = 64, word_weight: float = 0.35, name: str = "concept-v1"):
100        self.dim = dim
101        self.word_weight = word_weight
102        self.name = name
103
104    def _token_vector(self, tok: str) -> np.ndarray:
105        concept = WORD_TO_CONCEPT.get(tok)
106        word_vec = _unit_vector(f"{self.name}/word/{tok}", self.dim)
107        if concept is None:
108            return word_vec
109        concept_vec = _unit_vector(f"{self.name}/concept/{concept}", self.dim)
110        return concept_vec + self.word_weight * word_vec
111
112    def encode_one(self, text: str) -> np.ndarray:
113        toks = tokenize(text)
114        if not toks:
115            return np.zeros(self.dim)
116        v = np.sum([self._token_vector(t) for t in toks], axis=0)
117        return v / np.linalg.norm(v)
118
119    def encode(self, texts: str | Sequence[str] | Iterable[str]) -> np.ndarray:
120        """Embed one string (returns shape `(dim,)`) or many (returns `(n, dim)`)."""
121        if isinstance(texts, str):
122            return self.encode_one(texts)
123        return np.stack([self.encode_one(t) for t in texts])
CONCEPTS: dict[str, list[str]] = {'credential': ['password', 'passwords', 'passcode', 'credentials', 'login', 'sign-in', 'signin', 'logon'], 'reset': ['reset', 'recover', 'recovery', 'forgot', 'forgotten', 'unlock', 'locked', 'lockout'], 'mfa': ['2fa', 'mfa', 'authenticator', 'two-factor', 'otp', 'verification'], 'remote_access': ['vpn', 'remote', 'tunnel', 'anyconnect'], 'vehicle': ['car', 'automobile', 'vehicle', 'auto', 'mileage', 'driving'], 'money_back': ['reimburse', 'reimbursement', 'reimbursed', 'expense', 'expenses', 'refund', 'receipt', 'receipts'], 'time_off': ['pto', 'vacation', 'leave', 'holiday', 'holidays', 'time-off', 'sick'], 'computer': ['laptop', 'computer', 'notebook', 'macbook', 'device', 'hardware', 'monitor'], 'invoice': ['invoice', 'invoices', 'bill', 'billing', 'payment', 'payments', 'vendor', 'vendors'], 'travel': ['travel', 'trip', 'flight', 'flights', 'hotel', 'airfare', 'per-diem'], 'error': ['error', 'errors', 'failure', 'fail', 'failed', 'fails', 'crash', 'broken', 'issue'], 'email': ['email', 'emails', 'mail', 'inbox', 'outlook'], 'threat': ['phishing', 'suspicious', 'scam', 'malware', 'attack'], 'policy': ['policy', 'policies', 'rule', 'rules', 'requirement', 'requirements', 'guideline', 'guidelines'], 'new_hire': ['onboarding', 'onboard', 'hire', 'hires', 'orientation'], 'reconcile': ['reconcile', 'reconciliation', 'match', 'matching', 'mismatch', 'mismatches'], 'salary': ['salary', 'pay', 'payroll', 'paycheck', 'compensation', 'bonus'], 'printer': ['printer', 'printers', 'print', 'printing', 'toner']}
WORD_TO_CONCEPT: dict[str, str] = {'password': 'credential', 'passwords': 'credential', 'passcode': 'credential', 'credentials': 'credential', 'login': 'credential', 'sign-in': 'credential', 'signin': 'credential', 'logon': 'credential', 'reset': 'reset', 'recover': 'reset', 'recovery': 'reset', 'forgot': 'reset', 'forgotten': 'reset', 'unlock': 'reset', 'locked': 'reset', 'lockout': 'reset', '2fa': 'mfa', 'mfa': 'mfa', 'authenticator': 'mfa', 'two-factor': 'mfa', 'otp': 'mfa', 'verification': 'mfa', 'vpn': 'remote_access', 'remote': 'remote_access', 'tunnel': 'remote_access', 'anyconnect': 'remote_access', 'car': 'vehicle', 'automobile': 'vehicle', 'vehicle': 'vehicle', 'auto': 'vehicle', 'mileage': 'vehicle', 'driving': 'vehicle', 'reimburse': 'money_back', 'reimbursement': 'money_back', 'reimbursed': 'money_back', 'expense': 'money_back', 'expenses': 'money_back', 'refund': 'money_back', 'receipt': 'money_back', 'receipts': 'money_back', 'pto': 'time_off', 'vacation': 'time_off', 'leave': 'time_off', 'holiday': 'time_off', 'holidays': 'time_off', 'time-off': 'time_off', 'sick': 'time_off', 'laptop': 'computer', 'computer': 'computer', 'notebook': 'computer', 'macbook': 'computer', 'device': 'computer', 'hardware': 'computer', 'monitor': 'computer', 'invoice': 'invoice', 'invoices': 'invoice', 'bill': 'invoice', 'billing': 'invoice', 'payment': 'invoice', 'payments': 'invoice', 'vendor': 'invoice', 'vendors': 'invoice', 'travel': 'travel', 'trip': 'travel', 'flight': 'travel', 'flights': 'travel', 'hotel': 'travel', 'airfare': 'travel', 'per-diem': 'travel', 'error': 'error', 'errors': 'error', 'failure': 'error', 'fail': 'error', 'failed': 'error', 'fails': 'error', 'crash': 'error', 'broken': 'error', 'issue': 'error', 'email': 'email', 'emails': 'email', 'mail': 'email', 'inbox': 'email', 'outlook': 'email', 'phishing': 'threat', 'suspicious': 'threat', 'scam': 'threat', 'malware': 'threat', 'attack': 'threat', 'policy': 'policy', 'policies': 'policy', 'rule': 'policy', 'rules': 'policy', 'requirement': 'policy', 'requirements': 'policy', 'guideline': 'policy', 'guidelines': 'policy', 'onboarding': 'new_hire', 'onboard': 'new_hire', 'hire': 'new_hire', 'hires': 'new_hire', 'orientation': 'new_hire', 'reconcile': 'reconcile', 'reconciliation': 'reconcile', 'match': 'reconcile', 'matching': 'reconcile', 'mismatch': 'reconcile', 'mismatches': 'reconcile', 'salary': 'salary', 'pay': 'salary', 'payroll': 'salary', 'paycheck': 'salary', 'compensation': 'salary', 'bonus': 'salary', 'printer': 'printer', 'printers': 'printer', 'print': 'printer', 'printing': 'printer', 'toner': 'printer'}
class ConceptEmbedder: on GitHub
 87class ConceptEmbedder:
 88    """Deterministic toy bi-encoder. `encode(texts)` returns L2-normalized rows.
 89
 90    Args:
 91        dim: output dimensionality (real models use 384 to 3072).
 92        word_weight: how much each word's own identity counts relative to its
 93            concept. Lower means synonyms are closer together.
 94        name: a "model version" string. Two embedders with different names
 95            produce incompatible vector spaces, which lets
 96            `primer.ml.embeddings.operations` demonstrate why switching
 97            models means re-embedding everything.
 98    """
 99
100    def __init__(self, dim: int = 64, word_weight: float = 0.35, name: str = "concept-v1"):
101        self.dim = dim
102        self.word_weight = word_weight
103        self.name = name
104
105    def _token_vector(self, tok: str) -> np.ndarray:
106        concept = WORD_TO_CONCEPT.get(tok)
107        word_vec = _unit_vector(f"{self.name}/word/{tok}", self.dim)
108        if concept is None:
109            return word_vec
110        concept_vec = _unit_vector(f"{self.name}/concept/{concept}", self.dim)
111        return concept_vec + self.word_weight * word_vec
112
113    def encode_one(self, text: str) -> np.ndarray:
114        toks = tokenize(text)
115        if not toks:
116            return np.zeros(self.dim)
117        v = np.sum([self._token_vector(t) for t in toks], axis=0)
118        return v / np.linalg.norm(v)
119
120    def encode(self, texts: str | Sequence[str] | Iterable[str]) -> np.ndarray:
121        """Embed one string (returns shape `(dim,)`) or many (returns `(n, dim)`)."""
122        if isinstance(texts, str):
123            return self.encode_one(texts)
124        return np.stack([self.encode_one(t) for t in texts])

Deterministic toy bi-encoder. encode(texts) returns L2-normalized rows.

Arguments:

  • dim: output dimensionality (real models use 384 to 3072).
  • word_weight: how much each word's own identity counts relative to its concept. Lower means synonyms are closer together.
  • name: a "model version" string. Two embedders with different names produce incompatible vector spaces, which lets primer.ml.embeddings.operations demonstrate why switching models means re-embedding everything.
ConceptEmbedder(dim: int = 64, word_weight: float = 0.35, name: str = 'concept-v1') on GitHub
100    def __init__(self, dim: int = 64, word_weight: float = 0.35, name: str = "concept-v1"):
101        self.dim = dim
102        self.word_weight = word_weight
103        self.name = name
dim
name
def encode_one(self, text: str) -> numpy.ndarray: on GitHub
113    def encode_one(self, text: str) -> np.ndarray:
114        toks = tokenize(text)
115        if not toks:
116            return np.zeros(self.dim)
117        v = np.sum([self._token_vector(t) for t in toks], axis=0)
118        return v / np.linalg.norm(v)
def encode(self, texts: Union[str, Sequence[str], Iterable[str]]) -> numpy.ndarray: on GitHub
120    def encode(self, texts: str | Sequence[str] | Iterable[str]) -> np.ndarray:
121        """Embed one string (returns shape `(dim,)`) or many (returns `(n, dim)`)."""
122        if isinstance(texts, str):
123            return self.encode_one(texts)
124        return np.stack([self.encode_one(t) for t in texts])

Embed one string (returns shape (dim,)) or many (returns (n, dim)).