primer.common.embedder
ConceptEmbedder: a deterministic, dependency-free stand-in for a real
sentence-embedding model.
How it works (and why it's a fair teaching model):
- Tokenize the text (
primer.common.text.tokenize). - For each token, look it up in a tiny synonym lexicon (
CONCEPTS). If the word belongs to a concept (e.g. "car" and "automobile" both belong tovehicle), add that concept's fixed random vector. Also add a smaller word-specific vector so synonyms are close but not identical. - Words outside the lexicon (including IDs like "err-4012") get only their own random vector. They match themselves and nothing else.
- Mean-pool (sum) the token vectors and L2-normalize, exactly like a real bi-encoder does with mean pooling.
That reproduces the two behaviors that matter most in practice:
- Dense retrieval finds meaning. "automobile reimbursement" is close to "car expense" even with zero shared words.
- Dense retrieval blurs exact tokens. An error code is just one of many
pooled token vectors, so it's a weak signal next to the concept words.
That's why hybrid search (BM25 + dense) wins on enterprise data. See
primer.ml.embeddings.retrieval.
Random high-dimensional vectors are nearly orthogonal (their cosine is about 0 plus or minus 1/sqrt(dim)), which is what makes "unrelated" words unrelated here.
Further reading:
- Sentence-BERT (mean pooling, bi-encoders): https://arxiv.org/abs/1908.10084
- sentence-transformers docs: https://www.sbert.net/
- Random projections / near-orthogonality: https://en.wikipedia.org/wiki/Johnson%E2%80%93Lindenstrauss_lemma
1""" 2`ConceptEmbedder`: a deterministic, dependency-free stand-in for a real 3sentence-embedding model. 4 5How it works (and why it's a fair teaching model): 6 71. Tokenize the text (`primer.common.text.tokenize`). 82. For each token, look it up in a tiny synonym lexicon (`CONCEPTS`). If the 9 word belongs to a concept (e.g. "car" and "automobile" both belong to 10 `vehicle`), add that concept's fixed random vector. Also add a smaller 11 word-specific vector so synonyms are close but not identical. 123. Words outside the lexicon (including IDs like "err-4012") get only their 13 own random vector. They match themselves and nothing else. 144. Mean-pool (sum) the token vectors and L2-normalize, exactly like a real 15 bi-encoder does with mean pooling. 16 17That reproduces the two behaviors that matter most in practice: 18 19* **Dense retrieval finds meaning.** "automobile reimbursement" is close to 20 "car expense" even with zero shared words. 21* **Dense retrieval blurs exact tokens.** An error code is just one of many 22 pooled token vectors, so it's a weak signal next to the concept words. 23 That's why hybrid search (BM25 + dense) wins on enterprise data. See 24 `primer.ml.embeddings.retrieval`. 25 26Random high-dimensional vectors are nearly orthogonal (their cosine is 27about 0 plus or minus 1/sqrt(dim)), which is what makes "unrelated" words 28unrelated here. 29 30Further reading: 31- Sentence-BERT (mean pooling, bi-encoders): https://arxiv.org/abs/1908.10084 32- sentence-transformers docs: https://www.sbert.net/ 33- Random projections / near-orthogonality: https://en.wikipedia.org/wiki/Johnson%E2%80%93Lindenstrauss_lemma 34""" 35 36from __future__ import annotations 37 38import hashlib 39from functools import lru_cache 40from typing import Iterable, Sequence 41 42import numpy as np 43 44from primer.common.text import tokenize 45 46# concept name -> words that express it. Keep it small and obvious. 47CONCEPTS: dict[str, list[str]] = { 48 "credential": ["password", "passwords", "passcode", "credentials", "login", "sign-in", "signin", "logon"], 49 "reset": ["reset", "recover", "recovery", "forgot", "forgotten", "unlock", "locked", "lockout"], 50 "mfa": ["2fa", "mfa", "authenticator", "two-factor", "otp", "verification"], 51 "remote_access": ["vpn", "remote", "tunnel", "anyconnect"], 52 "vehicle": ["car", "automobile", "vehicle", "auto", "mileage", "driving"], 53 "money_back": ["reimburse", "reimbursement", "reimbursed", "expense", "expenses", "refund", "receipt", "receipts"], 54 "time_off": ["pto", "vacation", "leave", "holiday", "holidays", "time-off", "sick"], 55 "computer": ["laptop", "computer", "notebook", "macbook", "device", "hardware", "monitor"], 56 "invoice": ["invoice", "invoices", "bill", "billing", "payment", "payments", "vendor", "vendors"], 57 "travel": ["travel", "trip", "flight", "flights", "hotel", "airfare", "per-diem"], 58 "error": ["error", "errors", "failure", "fail", "failed", "fails", "crash", "broken", "issue"], 59 "email": ["email", "emails", "mail", "inbox", "outlook"], 60 "threat": ["phishing", "suspicious", "scam", "malware", "attack"], 61 "policy": ["policy", "policies", "rule", "rules", "requirement", "requirements", "guideline", "guidelines"], 62 "new_hire": ["onboarding", "onboard", "hire", "hires", "orientation"], 63 "reconcile": ["reconcile", "reconciliation", "match", "matching", "mismatch", "mismatches"], 64 "salary": ["salary", "pay", "payroll", "paycheck", "compensation", "bonus"], 65 "printer": ["printer", "printers", "print", "printing", "toner"], 66} 67 68# Reverse index: word -> concept. 69WORD_TO_CONCEPT: dict[str, str] = {w: c for c, ws in CONCEPTS.items() for w in ws} 70 71 72def _seed(key: str) -> int: 73 # Python's built-in hash() is salted per process; md5 is stable across runs. 74 return int(hashlib.md5(key.encode()).hexdigest()[:16], 16) 75 76 77@lru_cache(maxsize=None) 78def _unit_vector(key: str, dim: int) -> np.ndarray: 79 """A fixed pseudo-random unit vector for `key`. Same key -> same vector.""" 80 v = np.random.default_rng(_seed(key)).standard_normal(dim) 81 v /= np.linalg.norm(v) 82 v.setflags(write=False) # cached arrays are shared; make them read-only 83 return v 84 85 86class ConceptEmbedder: 87 """Deterministic toy bi-encoder. `encode(texts)` returns L2-normalized rows. 88 89 Args: 90 dim: output dimensionality (real models use 384 to 3072). 91 word_weight: how much each word's own identity counts relative to its 92 concept. Lower means synonyms are closer together. 93 name: a "model version" string. Two embedders with different names 94 produce incompatible vector spaces, which lets 95 `primer.ml.embeddings.operations` demonstrate why switching 96 models means re-embedding everything. 97 """ 98 99 def __init__(self, dim: int = 64, word_weight: float = 0.35, name: str = "concept-v1"): 100 self.dim = dim 101 self.word_weight = word_weight 102 self.name = name 103 104 def _token_vector(self, tok: str) -> np.ndarray: 105 concept = WORD_TO_CONCEPT.get(tok) 106 word_vec = _unit_vector(f"{self.name}/word/{tok}", self.dim) 107 if concept is None: 108 return word_vec 109 concept_vec = _unit_vector(f"{self.name}/concept/{concept}", self.dim) 110 return concept_vec + self.word_weight * word_vec 111 112 def encode_one(self, text: str) -> np.ndarray: 113 toks = tokenize(text) 114 if not toks: 115 return np.zeros(self.dim) 116 v = np.sum([self._token_vector(t) for t in toks], axis=0) 117 return v / np.linalg.norm(v) 118 119 def encode(self, texts: str | Sequence[str] | Iterable[str]) -> np.ndarray: 120 """Embed one string (returns shape `(dim,)`) or many (returns `(n, dim)`).""" 121 if isinstance(texts, str): 122 return self.encode_one(texts) 123 return np.stack([self.encode_one(t) for t in texts])
CONCEPTS: dict[str, list[str]] =
{'credential': ['password', 'passwords', 'passcode', 'credentials', 'login', 'sign-in', 'signin', 'logon'], 'reset': ['reset', 'recover', 'recovery', 'forgot', 'forgotten', 'unlock', 'locked', 'lockout'], 'mfa': ['2fa', 'mfa', 'authenticator', 'two-factor', 'otp', 'verification'], 'remote_access': ['vpn', 'remote', 'tunnel', 'anyconnect'], 'vehicle': ['car', 'automobile', 'vehicle', 'auto', 'mileage', 'driving'], 'money_back': ['reimburse', 'reimbursement', 'reimbursed', 'expense', 'expenses', 'refund', 'receipt', 'receipts'], 'time_off': ['pto', 'vacation', 'leave', 'holiday', 'holidays', 'time-off', 'sick'], 'computer': ['laptop', 'computer', 'notebook', 'macbook', 'device', 'hardware', 'monitor'], 'invoice': ['invoice', 'invoices', 'bill', 'billing', 'payment', 'payments', 'vendor', 'vendors'], 'travel': ['travel', 'trip', 'flight', 'flights', 'hotel', 'airfare', 'per-diem'], 'error': ['error', 'errors', 'failure', 'fail', 'failed', 'fails', 'crash', 'broken', 'issue'], 'email': ['email', 'emails', 'mail', 'inbox', 'outlook'], 'threat': ['phishing', 'suspicious', 'scam', 'malware', 'attack'], 'policy': ['policy', 'policies', 'rule', 'rules', 'requirement', 'requirements', 'guideline', 'guidelines'], 'new_hire': ['onboarding', 'onboard', 'hire', 'hires', 'orientation'], 'reconcile': ['reconcile', 'reconciliation', 'match', 'matching', 'mismatch', 'mismatches'], 'salary': ['salary', 'pay', 'payroll', 'paycheck', 'compensation', 'bonus'], 'printer': ['printer', 'printers', 'print', 'printing', 'toner']}
WORD_TO_CONCEPT: dict[str, str] =
{'password': 'credential', 'passwords': 'credential', 'passcode': 'credential', 'credentials': 'credential', 'login': 'credential', 'sign-in': 'credential', 'signin': 'credential', 'logon': 'credential', 'reset': 'reset', 'recover': 'reset', 'recovery': 'reset', 'forgot': 'reset', 'forgotten': 'reset', 'unlock': 'reset', 'locked': 'reset', 'lockout': 'reset', '2fa': 'mfa', 'mfa': 'mfa', 'authenticator': 'mfa', 'two-factor': 'mfa', 'otp': 'mfa', 'verification': 'mfa', 'vpn': 'remote_access', 'remote': 'remote_access', 'tunnel': 'remote_access', 'anyconnect': 'remote_access', 'car': 'vehicle', 'automobile': 'vehicle', 'vehicle': 'vehicle', 'auto': 'vehicle', 'mileage': 'vehicle', 'driving': 'vehicle', 'reimburse': 'money_back', 'reimbursement': 'money_back', 'reimbursed': 'money_back', 'expense': 'money_back', 'expenses': 'money_back', 'refund': 'money_back', 'receipt': 'money_back', 'receipts': 'money_back', 'pto': 'time_off', 'vacation': 'time_off', 'leave': 'time_off', 'holiday': 'time_off', 'holidays': 'time_off', 'time-off': 'time_off', 'sick': 'time_off', 'laptop': 'computer', 'computer': 'computer', 'notebook': 'computer', 'macbook': 'computer', 'device': 'computer', 'hardware': 'computer', 'monitor': 'computer', 'invoice': 'invoice', 'invoices': 'invoice', 'bill': 'invoice', 'billing': 'invoice', 'payment': 'invoice', 'payments': 'invoice', 'vendor': 'invoice', 'vendors': 'invoice', 'travel': 'travel', 'trip': 'travel', 'flight': 'travel', 'flights': 'travel', 'hotel': 'travel', 'airfare': 'travel', 'per-diem': 'travel', 'error': 'error', 'errors': 'error', 'failure': 'error', 'fail': 'error', 'failed': 'error', 'fails': 'error', 'crash': 'error', 'broken': 'error', 'issue': 'error', 'email': 'email', 'emails': 'email', 'mail': 'email', 'inbox': 'email', 'outlook': 'email', 'phishing': 'threat', 'suspicious': 'threat', 'scam': 'threat', 'malware': 'threat', 'attack': 'threat', 'policy': 'policy', 'policies': 'policy', 'rule': 'policy', 'rules': 'policy', 'requirement': 'policy', 'requirements': 'policy', 'guideline': 'policy', 'guidelines': 'policy', 'onboarding': 'new_hire', 'onboard': 'new_hire', 'hire': 'new_hire', 'hires': 'new_hire', 'orientation': 'new_hire', 'reconcile': 'reconcile', 'reconciliation': 'reconcile', 'match': 'reconcile', 'matching': 'reconcile', 'mismatch': 'reconcile', 'mismatches': 'reconcile', 'salary': 'salary', 'pay': 'salary', 'payroll': 'salary', 'paycheck': 'salary', 'compensation': 'salary', 'bonus': 'salary', 'printer': 'printer', 'printers': 'printer', 'print': 'printer', 'printing': 'printer', 'toner': 'printer'}
87class ConceptEmbedder: 88 """Deterministic toy bi-encoder. `encode(texts)` returns L2-normalized rows. 89 90 Args: 91 dim: output dimensionality (real models use 384 to 3072). 92 word_weight: how much each word's own identity counts relative to its 93 concept. Lower means synonyms are closer together. 94 name: a "model version" string. Two embedders with different names 95 produce incompatible vector spaces, which lets 96 `primer.ml.embeddings.operations` demonstrate why switching 97 models means re-embedding everything. 98 """ 99 100 def __init__(self, dim: int = 64, word_weight: float = 0.35, name: str = "concept-v1"): 101 self.dim = dim 102 self.word_weight = word_weight 103 self.name = name 104 105 def _token_vector(self, tok: str) -> np.ndarray: 106 concept = WORD_TO_CONCEPT.get(tok) 107 word_vec = _unit_vector(f"{self.name}/word/{tok}", self.dim) 108 if concept is None: 109 return word_vec 110 concept_vec = _unit_vector(f"{self.name}/concept/{concept}", self.dim) 111 return concept_vec + self.word_weight * word_vec 112 113 def encode_one(self, text: str) -> np.ndarray: 114 toks = tokenize(text) 115 if not toks: 116 return np.zeros(self.dim) 117 v = np.sum([self._token_vector(t) for t in toks], axis=0) 118 return v / np.linalg.norm(v) 119 120 def encode(self, texts: str | Sequence[str] | Iterable[str]) -> np.ndarray: 121 """Embed one string (returns shape `(dim,)`) or many (returns `(n, dim)`).""" 122 if isinstance(texts, str): 123 return self.encode_one(texts) 124 return np.stack([self.encode_one(t) for t in texts])
Deterministic toy bi-encoder. encode(texts) returns L2-normalized rows.
Arguments:
- dim: output dimensionality (real models use 384 to 3072).
- word_weight: how much each word's own identity counts relative to its concept. Lower means synonyms are closer together.
- name: a "model version" string. Two embedders with different names
produce incompatible vector spaces, which lets
primer.ml.embeddings.operationsdemonstrate why switching models means re-embedding everything.
120 def encode(self, texts: str | Sequence[str] | Iterable[str]) -> np.ndarray: 121 """Embed one string (returns shape `(dim,)`) or many (returns `(n, dim)`).""" 122 if isinstance(texts, str): 123 return self.encode_one(texts) 124 return np.stack([self.encode_one(t) for t in texts])
Embed one string (returns shape (dim,)) or many (returns (n, dim)).