Pack format
The normative format of the data packs, vectors and shards, for ports to any language.
On this page
This document is normative. A client in any language (TypeScript, Swift, Kotlin) that follows it
gets the same search results as packages/core. The words MUST and SHOULD have their RFC 2119
meaning.
A pack version (e.g. 0.1.0) is a directory of immutable files:
/v1/pack/<packVersion>/
manifest.json versions, hashes, sizes
pack.en.json Tier 0 core, English (always load first; ≤ 200 KB gz)
pack.en.ext.json Tier 0 extension, English (load when idle)
pack.<locale>.json / .ext.json core / extension of every other locale (zh, hi, es, ar, fr,
bn, pt, ru, id, tr; load next to English for that UI locale)
vectors.<model>.<dims>.bin Tier 1 emoji vectors, one file per model × dims (English documents)
vectors.<model>.<dims>.<locale>.bin Tier 1 emoji vectors of one locale's documents (optional, §5)Files never change after publication. A change produces a new pack version. Serve them with
Cache-Control: public, max-age=31536000, immutable.
1. manifest.json
{
"format": "emojisense-manifest",
"formatVersion": 1,
"packVersion": "0.1.0",
"emojiVersion": "17.0",
"emojiCount": 1914,
"source": { "emojibaseVersion": "17.0.0", "cldrVersion": "48.2.0" },
"coreAliases": { "en": 16, "hi": 14 },
"files": {
"pack.en.json": { "sha256": "…", "bytes": 1234, "gzipBytes": 456, "locale": "en" },
"vectors.bge-m3.1024.bin": {
"sha256": "…", "bytes": 2229760, "gzipBytes": 1751207,
"model": "@cf/baai/bge-m3", "dims": 1024,
"queryTemplate": "{q}"
},
"vectors.bge-m3.1024.es.bin": {
"sha256": "…", "bytes": 2229760, "gzipBytes": 1762125, "locale": "es",
"model": "@cf/baai/bge-m3", "dims": 1024,
"queryTemplate": "{q}"
}
}
}- Clients SHOULD verify
sha256(lowercase hex of the raw file bytes) before they cache a file. coreAliasesis informational: how many aliases per emoji each locale's core part keeps (§2).queryTemplateis the exact string to embed for a query.{q}is replaced by the query's embedding text (§3, "Embedding text"). Only needed by clients that embed queries themselves.- A vector file with a
localekey holds that locale's document vectors (§5).
2. pack..json
{
"format": "emojisense-pack",
"formatVersion": 1,
"packVersion": "0.1.0",
"locale": "en",
"emojiVersion": "17.0",
"groups": ["smileys-emotion", "people-body", "…"],
"weights": { "alias": 0.8 },
"emoji": [
["🦖", "1F996", 2, 5, 0, "T-Rex", "t rex", "dinosaur|rex|t|t rex|tyrannosaurus", "jurassic park|dino|…", "dinasour|…", "…"]
]
}A client MUST reject a pack whose format differs or whose formatVersion it does not support.
Core and extension parts
Each locale ships in two parts with the same row layout and row order:
| Part | File | part key |
Holds |
|---|---|---|---|
| core | pack.<locale>.json |
absent or "core" |
label, shortcodes, keywords, the first N aliases (N ≤ pack.config.json → initialAliases, lowered per locale until the part is ≤ 200 KB gz; manifest coreAliases) |
| ext | pack.<locale>.ext.json |
"ext" |
the remaining aliases, all typos, all low-confidence phrases. label, shortcode and keyword are empty. |
Render with the core parts. Load the ext parts when the device is idle and rebuild the index with
all parts. Index order: every core part first (English first), then the ext parts. All parts of
one locale count as that locale for the preferred-locale factor (§4). An empty label MUST NOT
replace a label from another part.
A third part, "custom", holds an app's own emoji. It is not a file of a pack version: the API
builds it per app (§8).
Rows
Each element of emoji is an array with 11 positions. Row order is the display order (CLDR
order). Every locale pack of one pack version has the same rows in the same order.
| # | Name | Type | Meaning |
|---|---|---|---|
| 0 | emoji | string | The emoji, fully qualified (with U+FE0F where CLDR has it). |
| 1 | hexcode | string | Emojibase hexcode of the base emoji, e.g. 1F44D. This is the stable id. |
| 2 | group | int | Index into groups. |
| 3 | version | number | Emoji version that introduced it (e.g. 15.1). Hide rows the OS cannot render. |
| 4 | skins | 0 | 1 | 1 = supports skin tones. Variant = hexcode + -1F3FB … -1F3FF. |
| 5 | label | string | Display label in this locale. Not normalized. |
| 6 | shortcode | phrases | Shortcodes (GitHub, Slack-style, Emojibase), normalized. |
| 7 | keyword | phrases | CLDR keywords for this locale. |
| 8 | alias | phrases | Generated aliases: synonyms, slang, pop culture, dev, intent. Strongest first. |
| 9 | typo | phrases | Common misspellings. |
| 10 | low | phrases | Low-confidence or demoted aliases. |
phrases = phrases joined with |. Each phrase is already normalized (§3), so it contains
only [a-z0-9+ ] after folding plus letters of other scripts. It never contains |. An empty
string means no phrases.
Search fields, strongest first, with default weights (a pack MAY override them in weights):
| Field | Source | Weight |
|---|---|---|
| name | normalize(label) (derived, not stored) |
1.00 |
| shortcode | row[6] | 0.95 |
| keyword | row[7] | 0.85 |
| alias | row[8] | 0.80 |
| typo | row[9] | 0.75 |
| low | row[10] | 0.55 |
When the same phrase occurs twice for one emoji, only the strongest field counts.
3. Normalization
Apply to every query, and to label to get the name field. In order:
- Unicode NFKC.
- Replace every code point with property
Extended_Pictographic,Emoji_ModifierorRegional_Indicator, and U+E0020–U+E007F, with a space. - Remove U+200D, U+FE0E, U+FE0F and U+20E3.
- Lowercase (locale-independent Unicode default case mapping).
- Unicode NFD, then remove only the optional marks: U+0300–U+036F (Latin, Greek and Cyrillic diacritics), U+064B–U+065F and U+0670 (Arabic harakat), U+0640 (tatweel), and U+0591–U+05C7 (Hebrew points). Keep all other marks: in Devanagari, Bengali, Thai or Japanese they are part of the spelling.
- Replace
ı→i,đ→d,ł→l,ø→o,ß→ss. - Unicode NFC (recomposes Hangul and kana).
- Remove the apostrophes
'’`´. - Replace each run of characters that are not a letter (
L), a mark (M), a number (N) or+with one space. - Replace each
+that is not followed by a digit with a space. - Collapse whitespace runs to one space and trim.
- Truncate to 64 UTF-16 code units, then trim the end.
Scripts without spaces between words (Chinese, Japanese, Thai) form one token per run. Prefix matching still completes them while the user types. A query run that is not one indexed token is split at search time (§4, "Unspaced scripts").
Examples: "İYİ Kİ DOĞDUN" → "iyi ki dogdun", "¡Feliz cumpleaños!" → "feliz cumpleanos",
"Ёлка" → "елка", "مَرْحَبًا" → "مرحبا", "नमस्ते" → "नमस्ते" (unchanged), "i'm exhausted" → "im exhausted",
":rocket:" → "rocket", "+1" → "+1", "🚀 launch 👍🏽" → "launch".
Tokens are the result split on single spaces.
Embedding text. The semantic tier embeds a lighter form of the query, because the embedding
model reads accents and punctuation (folding them cost about 3 points of semantic recall@5):
Unicode NFKC, lowercase, NFKC again, each run of \p{Cc}, \p{Z} or U+FEFF to one space, trim,
then truncate to 64 UTF-16 code units without splitting a surrogate pair. Accents, punctuation
and emoji stay: " Doğum GÜNÜ!! " → "doğum günü!!". Its normalized form (steps 1–12) is the
normalized query. Reference: embeddingText in packages/core/src/normalize.ts.
4. Tier 0 search (reference algorithm)
The reference implementation is packages/core/src/engine.ts. A port SHOULD match it, so that
the shared eval set gives the same numbers on every platform.
Index. For every row of every loaded pack, collect phrases per field. Build a sorted
vocabulary of all tokens, postings token → phrases, and for each token
idf = ln(1 + E / df), where E is the emoji count and df is the number of distinct emoji
that have the token in any phrase.
Query. Normalize and tokenize the query, keeping at most 8 tokens.
Unspaced scripts. A query token that holds a code point in U+0E00–0EFF (Thai, Lao),
U+1000–109F (Myanmar), U+1780–17FF (Khmer), U+3040–30FF (kana), U+3400–4DBF, U+4E00–9FFF,
U+F900–FAFF or U+20000–3134F (Han) is split when it is not a vocabulary token, unless it is the
last token while typing and a longer vocabulary token starts with it. Split from the left: at each
code point take the longest vocabulary token (at most 16 code points) that starts there. Code
points where no vocabulary token starts form one unknown piece together with the unknown code
points next to them. The pieces replace the token, in order; keep at most 8 tokens again. The
pieces are the query's tokens. Example, with 生日快乐 indexed: 今天生日快乐 → 今天 · 生日快乐.
For each query token, find candidate vocabulary tokens with a match quality:
| Match | Quality |
|---|---|
| exact | 1.0 |
| prefix (only the last token, only while typing, i.e. the raw query does not end with whitespace) | 0.6 + 0.35 × len(query token) / len(vocab token) |
final repeated letter removed (upp → up), only if there is no exact match |
0.85 |
| optimal-string-alignment distance 1 (token length 4–7) or ≤ 2 (length ≥ 8), only if there is no exact match | 0.8 (d = 1), 0.65 (d = 2) |
The weight of query token i is the idf of its best candidate, or the largest idf in the index
when there is no candidate. For the stopwords listed in engine.ts, the weight is capped at 0.3.
Phrase score. coverage = Σ quality_i × weight_i / Σ weight_i, with quality_i the best
quality of a candidate of token i inside this phrase. Phrases with coverage < 0.34 are
ignored. Then:
score = fieldWeight × coverage × (0.6 + 0.4 × min(1, matchedTokens / phraseTokens))
× (exactPhrase ? (query tokens ≥ 2 ? 1.1 : 1) : 0.9) × (phrase in preferred-locale pack ? 1 : 0.92)exactPhrase = every query token matched exactly and the phrase has as many tokens as the query.
The preferred locale is the query's locale, or the locale of the first loaded pack when the
query has none. A phrase is in a preferred-locale pack when any pack of that locale (core or ext)
contains it for this emoji.
Emoji score. The best phrase score, plus 0.02 for every other matching phrase that is in a preferred-locale pack (at most +0.06), capped at 1. Phrases of other locales never add this bonus, so many loaded languages that share a loanword ("halloween") cannot lift every emoji to the cap.
Whole query before a partial match. The bonus breaks near-ties only. For each emoji whose
best phrase is not an exactPhrase match, let W be the lowest emoji score among the emoji
whose best phrase is an exactPhrase match in a preferred-locale pack and has a higher phrase
score (before the bonus). When there is one, the emoji scores at most W − 0.01. Without this,
en "ship it" gave 🚢 (name ship, the stopword uncovered, plus +0.06 from its other ship
phrases) before 🚀 (alias ship it).
Preferred exact match first. Let P be the highest emoji score among the emoji that have
an exactPhrase match in the name, shortcode, keyword or alias field of a
preferred-locale pack. When there is one, every emoji without such a match whose best phrase is
an exactPhrase match in the name or shortcode field of another pack scores at most
P − 0.01. These two fields outweigh a preferred keyword or alias even after the foreign
factor, so without this fr "foot" gave 🦶 (English name foot) before ⚽ (French alias foot).
Sort by score (descending), then by row order. confidence = the top score.
5. Vectors (vectors.<model>.<dims>.bin, "ESVEC1")
Little-endian. All offsets are in bytes from the file start.
| Offset | Size | Field |
|---|---|---|
| 0 | 8 | magic "ESVEC1\0\0" (ASCII) |
| 8 | 4 | u32 count (rows) |
| 12 | 4 | u32 dims (multiple of 8) |
| 16 | 4 | u32 modelLength (bytes) |
| 20 | 4 | u32 idsLength (bytes) |
| 24 | 8 | reserved, zero |
| 32 | modelLength | model id, UTF-8 (e.g. @cf/baai/bge-m3) |
| A = align4(32 + modelLength) | idsLength | hexcodes joined with \n, UTF-8; row order |
| S = align4(A + idsLength) | 4 × count | f32 scale per row |
| V = S + 4 × count | count × dims | i8 quantized components, row-major |
| V + count × dims | count × dims / 8 | sign bits, row-major; bit k of the row is byte (r·dims + k) >> 3, bit (r·dims + k) & 7; 1 = component > 0 |
- Component
dof rowr≈i8[r·dims + d] × scale[r]. Before quantization, each row was L2-normalized, so a dot product with a normalized query is cosine similarity. - Rows were embedded at the model's native dims, truncated to
dims(Matryoshka), and then L2-normalized again. A query vector MUST get the same treatment. - Queries MUST be embedded with the same model, using
queryTemplatefrom the manifest. A vector file must never be compared with vectors from another model or with other dims. - Sign bits allow a cheap Hamming-distance shortlist on weak devices before an int8 rerank.
Shared and locale files. Each model × dims has one shared file, embedded from English
documents. A multilingual model MAY also have one locale file per pack locale,
vectors.<model>.<dims>.<locale>.bin (e.g. vectors.bge-m3.1024.es.bin), embedded from that
locale's documents. A document is "<label>. <description> <CLDR keywords>, <first 40 aliases>"
in its language (packages/data/src/documents.ts). All files of one model × dims have the same
binary layout, model, dims and rows (hexcodes in pack order). In the manifest, a locale file
has a locale key; the shared file has none.
- A query of locale
Lscores each emoji by its best row over the shared file andL's file:max(cos(q, shared[e]), cos(q, L[e])). Without a file forL(English, or a locale that has none), the shared file alone gives the ranking. Reference:searchVectorSetsinpackages/core. - A client MAY load the shared file only. Its results stay valid; they are the English-document ranking.
6. Shards (layer 2: precomputed results)
Frequent queries that the on-device dictionary cannot answer get their semantic results precomputed nightly and published as static files:
/p/<packVersion>/index.json {"format":"emojisense-shards","formatVersion":1,"packVersion":"0.1.0",
"model":"bge-m3@1024","keys":["a","ab","b", … ,"th","the ", …]}
/p/<packVersion>/<key>.json {"key":"co","entries":{"congrats on the launch":[["🚀","1F680",0.81], …]}}keysare sorted. A query uses the longest key that is a prefix of the normalized query. Hot prefixes get longer keys (adaptive split), so each shard stays ≤ ~30 KB gz. File names areencodeURIComponent(key).entriesmaps a normalized query (§3) to semantic results[emoji, hexcode, score], best first. These are the same results the API returns withmode=semanticfor that model andlocale=en: shards are built from the shared vector file only (§5).- A client downloads
index.jsononce and each shard at most once per session, then answers locally. A query that is not in its shard goes to the API. - Shards are valid only for the
modelthey name. A new model or pack version publishes a new directory. - A client uses a shard only when
embeddingText(query)equalsnormalize(query)(§3). The API embeds the text as typed, accents and punctuation kept, so a query such as "doğum günü" or "i'm done!" goes to the API instead of taking the answer of its folded form. - The API Worker rebuilds the shards every night from the query counts (ARCHITECTURE.md, "Nightly
shard build") and serves them at the same URLs.
index.jsonand the key files therefore change under one pack version: they are cached for 1 hour (index.json) and 1 day (key files), neverimmutable. Key files of an older build hold valid answers for the same data; a key that is gone answers 404, and the client asks the API. - No key is ever
index: its file would replaceindex.json.
7. Versioning
formatVersionchanges only for breaking layout changes. Adding optional manifest keys or pack-level keys is not breaking. Adding a row position is breaking.packVersionis semver for the data. A patch has alias changes only. A minor adds emoji or locales. A major changes ids.
8. Custom packs (an app's own emoji)
GET /v1/custom-pack?key=…[&tenant=<externalId>] (docs/API.md) returns the custom emoji of the
key's app, plus those of one tenant, as a pack of formatVersion 1 with the same row layout:
{
"format": "emojisense-pack",
"formatVersion": 1,
"packVersion": "custom-3f9a0c1e",
"locale": "und",
"part": "custom",
"emojiVersion": "",
"groups": ["custom"],
"emoji": [
[":party_parrot:", "C-x7Kq2", 0, 0, 0, "party_parrot", "party parrot", "", "celebrate|dance", "", ""]
],
"images": { "C-x7Kq2": "https://api.emojisense.com/v1/custom/app_1/x7Kq2" }
}| Key or row position | Value |
|---|---|
part |
"custom" |
locale |
"und" (BCP 47 "undetermined"): custom emoji belong to no locale |
packVersion |
custom- + 8 hex characters, a hash of the rows and images. It changes when the set changes. |
images |
Image URL per hexcode. Images are immutable (docs/API.md, GET /v1/custom/:appId/:emojiId). |
| row[0] emoji | :shortcode:. Draw the image of images[row[1]] instead of text; use :shortcode: as its alt text. |
| row[1] hexcode | C-<emojiId>. The C- prefix never collides with an Emojibase hexcode. |
| row[2] group | 0 = groups[0] = "custom" |
| row[3] version | 0: every device can draw an image |
| row[4] skins | 0: skin tones do not apply |
| row[5] label | the shortcode without colons; its normalized form is the name field |
| row[6] shortcode | the normalized shortcode (party_parrot → party parrot) |
| row[7] keyword | empty |
| row[8] alias | the emoji's aliases, normalized (§3) |
| row[9], row[10] | empty |
A tenant emoji replaces an app-wide emoji with the same shortcode. Rows are sorted by shortcode.
Search. Load a custom pack after the locale packs and index all of them together:
- Primary pack: the first pack that is not custom. Custom rows are appended as their own entries (they are not merged into catalog rows) and take part in IDF like any other entry.
- Custom phrases count for every locale: the preferred-locale factor (§4) is always 1 for them.
- A matching custom row is a result with
source: "custom"and the extra fieldsimageUrl(fromimages) andshortcode. Catalog results keep their shape. - The engine's
localesdo not includeund.
The API merges the same custom matches into /v1/search and /v1/suggest-reactions, first. A
client that fuses its own results with the API's sees each custom emoji once (fusion is by id).
Clients without custom-pack support can ignore packs with part: "custom": their layout is valid
and their rows never match catalog hexcodes.
9. Culture files (the culture layer)
Editorial associations that add emoji next to the canonical answer, by culture, region and moment (docs/ARCHITECTURE.md, "Culture layer"). One small file per locale, built at deploy for the next 12 months; clients decide by their own day what is active, so it needs no daily rebuild:
/v1/culture/<packVersion>/culture.<locale>.json Cache-Control: public, max-age=3600 (not immutable)
/v1/culture/<packVersion>/index.json build date, window and per-locale sizes (informational){
"format": "emojisense-culture",
"formatVersion": 1,
"packVersion": "0.1.0",
"locale": "es",
"from": "2026-10-02",
"until": "2027-10-03",
"entries": [
{
"id": "goat-football",
"kind": "lasting",
"context": "El debate sobre el mejor futbolista de la historia",
"when": null,
"regions": ["*"],
"triggers": ["goat", "el goat", "el mejor de la historia"],
"emoji": [["🐐", "1F410", 0.7], ["⚽", "26BD", 0.6], ["🇦🇷", "1F1E6-1F1F7", 0.45]]
},
{
"id": "halloween",
"kind": "seasonal",
"context": "Halloween, 31 de octubre",
"when": { "from": "10-15", "to": "10-31", "recurs": "yearly" },
"regions": ["*"],
"triggers": ["halloween", "noche de brujas"],
"emoji": [["🎃", "1F383", 0.9], ["👻", "1F47B", 0.75]],
"featured": true
}
],
"relevantNow": []
}A client MUST reject a file whose format differs or whose formatVersion it does not support.
| Key | Meaning |
|---|---|
from, until |
Days the build covered: the file holds every lasting entry, plus the seasonal and event entries active on any day of [from, until]. A build covers 366 days (at least 12 months, also across a leap day), so every yearly entry is in the file. |
entries[].kind |
lasting, seasonal (a yearly window), event (one dated window, ≤ 60 days) or regional (a word whose main sense differs by region, always active; see step 3). A festival on a lunar calendar is one event entry per year (diwali-2026). |
entries[].context |
The reason, in this file's locale. Neutral, ≤ 90 characters. |
entries[].when |
null (always), { from: "MM-DD", to: "MM-DD", recurs: "yearly" } (may wrap the year end, e.g. 12-26 → 01-02) or { from: "YYYY-MM-DD", to: "YYYY-MM-DD" }. Days are inclusive and compared with the user's local calendar day (the search API: the request's UTC day). |
entries[].regions |
ISO 3166-1 alpha-2 codes, or ["*"]. Without a region from the app, only "*" entries apply. |
entries[].exceptRegions |
Optional, with regions: ["*"]: codes where the entry does not apply when the app names one of them. |
entries[].outranks |
regional entries only: hexcodes of the canonical top answers the regional sense may move to second place. |
entries[].triggers |
Normalized phrases (§3) that people of this locale type. |
entries[].emoji |
[emoji, hexcode, weight], strongest first; weight in (0, 1]. Base hexcodes only. |
entries[].featured |
May appear on an optional "relevant now" shelf (seasonal and event entries only). |
relevantNow |
Always [] in 12-month files (see below). In older files: ids of the featured entries active on from, in shelf order, for clients that do not evaluate windows. |
Applying it (reference: packages/core/src/culture.ts).
- Normalize the query (§3). An entry applies when its window is active today and its regions
match. A trigger matches when it equals the query (quality 1), or, while the user is typing
(no trailing space), when the query is a prefix of the trigger with ≥ 3 characters and at least
half its length (quality
0.6 + 0.4 × len(query) / len(trigger)). - Score each emoji
weight × quality; keep the best score per hexcode; take the best 5. Drop emoji the loaded packs do not have. - Insert them right after the canonical top result (after fusion with semantic results),
skipping the top result's own emoji; an emoji that is already lower in the list moves up.
Never put a culture emoji above the canonical top result, except when the canonical list is
empty, or for a regional sense: when the app names a region in a
regionalentry's scope, the normalized query equals one of its triggers (not a prefix) and the canonical top result is one of itsoutranks, its strongest emoji goes first and the canonical top result second (several qualify: the strongest wins). Cut the list to the requested limit. - Mark them
source: "culture"withcontextandcultureId. An option to turn the layer off (culture: false) MUST exist for reproducible ranking.
regional, exceptRegions and outranks came after the first files, under the same
formatVersion: 1. A client that does not know them treats a regional entry as a lasting one and
ignores exceptRegions: it adds the emoji after the top result (also in an excluded region), never
above it. That is safe under the add-never-replace rule.
The relevant now shelf lists the featured seasonal and event entries active today (in file order, which puts events first), taking one emoji from each entry in turn.
The source format (packages/data/culture/entries/<id>.json, one file per association with a
status, context and triggers per locale, and provenance) is described by
packages/data/culture/schema.json.
12-month files (2026-10-02), same formatVersion: 1. Files used to cover 14 days and needed a
daily rebuild. Now a build covers 366 days from the day before the deploy (UTC, so a device west
of UTC that is still on that day finds it), and the client checks every entry's when against its
own day at query time, as step 1 always required. The keys and their types did not change, so
old clients keep loading the files:
| Client | Behaviour with a 12-month file |
|---|---|
| Checks windows (every emojisense SDK so far, the search API) | Correct on every day of the 12 months, with no new download. |
Reads relevantNow instead of checking windows |
Gets [] and shows no shelf, never an out-of-date one. |
Ignores when (breaks step 1) |
Already wrong with 14-day files; now shows seasonal entries all year. Fix the client. |
The files change only at a deploy, so new or edited entries still need a sync and a deploy. A
file whose until has passed still works for lasting and yearly entries but misses later events.