Imported from pinecone-io/skills (
skills/pinecone-full-text-search/SKILL.md). Install upstream withnpx skills add pinecone-io/skills --skill pinecone-full-text-search. Copyright stays with the author.
Pinecone Full-Text Search
Requires
pineconePython SDK ≥ 10.0.0 (pip install pinecone>=10.0.0). The document-schema API graduated out ofpinecone.previewin 10.0.0 — it is now a first-class, SemVer-covered part of the SDK, reachable directly offpc(pc.indexes,pc.index(...)). If you land on this skill from an older habit of importingpinecone.preview, stop: that package is deleted outright in 10.0.0 (ModuleNotFoundError, no shim). The packaged helper script pinspinecone==10.0.0via PEP 723 inline metadata; if you're writing your own code against this skill, pin at least that version. The wire API version is2026-07.
Authoritative reference (last resort). If you hit a question this skill and its
references/*.mdfiles don't answer, the official Pinecone FTS docs are at https://docs.pinecone.io/guides/search/full-text-search. Prefer this skill's content for anything covered here — the docs may describe surfaces (e.g. classic vector API, or the olderpinecone.previewshape) that don't apply to the graduated document-schema path. Consult the link only when you're genuinely stuck.
Tell the user up front: "This skill ships a helper at
scripts/ingest.pythat handles bulk ingestion safely (batched upsert, error inspection, readiness polling). When we get to the ingest step, I'll use it." Surface this at the start of the conversation so the user knows the helper exists. Query construction is hand-writtendocuments.search(...)per the Querying section below — there is no query helper.
A workflow skill for building a Pinecone full-text-search index with the graduated document-schema API (pc.indexes, pc.index(name), API version 2026-07). Covers schema design (text, dense vector, sparse vector, filterable metadata), ingestion (including async indexing and polling), and query construction (text / query_string / dense_vector / sparse_vector scoring; $match_phrase / $match_all / $match_any text-match filters; $eq / $in / $gte / $exists / $and / $or / $not metadata filters).
Scope — this skill is for the document-schema FTS API only
This skill covers pc.indexes.create(..., schema=...), pc.index(name), idx.documents.upsert(...) / idx.documents.batch_upsert(...) / idx.documents.search(...). If you find yourself reaching for any of the following, stop — those are different Pinecone APIs and this skill's guidance and helpers won't apply:
- Classic vector / records API:
pc.Index(name),index.upsert(vectors=[...]),index.query(vector=..., sparse_vector=...),pc.create_index(dimension=..., metric=..., spec=ServerlessSpec(...)). This is the deprecated sugar path in 10.0.0 — it still runs, but it creates a schemaless index served by the vector data plane, addressing the vector by the reserved_valuesfield. It cannot holdfull_text_searchfields. - Integrated-embedding / records indexes:
pc.create_index_for_model(...)/pc.indexes.create_for_model(...)withembed={...}. Pinecone vectorizes text server-side, and the resultingsemantic_textfield is served by the records API (upsert_records/search_records), not the documents API. Different upsert/search shapes. Asemantic_textfield cannot be combined withfull_text_searchfields in the same index.
If the user already has a non-document-schema index, they can stand up a separate document-schema index alongside it — the two are independent — but you can't add FTS fields to a classic or integrated-embedding index after the fact, and a document-schema index only ever serves reads and writes through index.documents.* — never index.upsert / index.query / index.upsert_records (those calls are refused with "This index has a document schema, so writes must go through the documents API").
Querying — construct documents.search(...) calls
For any task that asks you to query an FTS index, you write a documents.search(...) call directly. The schema is authoritative — describe the index live before constructing the call so you know which fields are FTS-enabled, which are filterable, and which are vectors.
Workflow:
- Discover the schema. Call
pc.indexes.describe(<index>)and read theschema.fieldsdict. Each field's class indicates its type (StringField,FloatField,DenseVectorField, etc.); attributes tell you whether it's FTS-enabled (full_text_search), filterable, or carries adimension. Skip this step only if you've already seen the schema in this conversation. - Construct the call matching the rules below — one scoring type per request, hard requirements in
filter, ranking signals inscore_by,include_fieldsexplicit on every call. - Execute with
idx = pc.index(name=<index>); resp = idx.documents.search(...)and readresp.matches.
Canonical shapes:
# Pure BM25 keyword search
resp = idx.documents.search(
namespace="__default__",
top_k=10,
score_by=[{"type": "text", "field": "body", "query": "machine learning"}],
filter={"year": {"$gt": 2024}, "category": {"$eq": "ai"}}, # optional
include_fields=["*"], # always pass explicitly
)
# Hybrid: dense ranking with a lexical filter (one type in score_by + filter narrows)
resp = idx.documents.search(
namespace="__default__",
top_k=10,
score_by=[{"type": "dense_vector", "field": "embedding", "values": query_embedding}],
filter={"body": {"$match_all": "TensorFlow"}, "year": {"$gt": 2024}},
include_fields=["*"],
)
Key rules (the server enforces these; following them locally keeps the agent loop tight):
score_byis a list of clauses, but exactly one scoring type per request (server rejects mixed types). Multi-field BM25 is the one exception: multipletextclauses, or onequery_stringwithfields: [...]. To combine BM25 + dense signals, restrict the dense search with a text-match filter ($match_all/$match_phrase/$match_any); do NOT mix scoring types inscore_by.filterkeys are field names (must exist in schema, or be an auto-indexed metadata field from upserted documents — see Filterable metadata isn't declared in the schema below) OR logical operators ($and,$or,$not). Field values are operator dicts ({"$gt": 5}, NOT bare values).include_fieldsis required on every call. Pass["*"]for all stored fields,[]for ids+score only, or a list of names. Omitting it on some SDK/backend builds 400s.
Clause shapes (for score_by):
type |
Required keys | When to pick this |
|---|---|---|
text |
field (string FTS), query |
Open-ended keyword search; BM25 ranking on one field |
query_string |
query (Lucene), fields optional |
Lucene boost (^N), proximity (~N), cross-field boolean, phrase prefix |
dense_vector |
field (dense_vector), values (list of floats) |
Semantic / mood / topic ranking |
sparse_vector |
field (sparse_vector), sparse_values ({indices, values}) |
Custom sparse-encoder ranking |
text / dense_vector / sparse_vector use singular field. Only query_string accepts a fields array (and also accepts singular field as an alias). sparse_vector uses sparse_values (NOT values) — distinct from dense.
Filter operators by field type:
| Field type | Legal operators |
|---|---|
string with FTS |
$match_phrase, $match_all, $match_any |
| filterable metadata (string / auto-indexed) | $eq, $ne, $in, $nin, $exists |
string_list filterable (auto-indexed, not schema-declared) |
$in, $nin, $exists |
float filterable (auto-indexed, not schema-declared) |
$eq, $ne, $gt, $gte, $lt, $lte, $exists |
boolean filterable (auto-indexed, not schema-declared) |
$eq, $exists |
| logical wrappers | $and: [filters], $or: [filters], $not: filter |
Match shape on response:
for m in resp.matches:
m._id # document id
m._score # match score (NOT `score`)
m.to_dict() # full doc payload (when include_fields includes the field)
For deeper coverage — multi-field BM25, Lucene patterns, hybrid composition, RRF merges, common error symptoms — see references/querying.md. For schema field types and what they enable on the query side, see references/schema-design.md.
Ingesting — use the packaged helper
For any task that asks you to bulk-ingest a JSONL file into an existing FTS index, the canonical path is to invoke the bundled helper, NOT to hand-write a Python script. Do not read the script's source — everything you need is in this section.
The script does three things bare-LLM ingest code reliably skips, each of which corresponds to a silent production failure:
- Bulk-upserts in batches. No per-doc
upsertloops. - Inspects every batch result.
batch_upsertreturns 202 even when individual documents fail; the failures live inresult.errors/result.has_errors. Without inspection, "100 docs ingested" silently becomes "73 docs ingested + 27 lost." - Polls until searchable. After upsert, Pinecone is still building the inverted index. A
documents.searchcall during that window returns empty. Without the poll, the user debugs their query code for an hour without finding the indexing race.
You provide a prepared, schema-conformant JSONL file and the index name; the script does the rest. Schema validation is upstream concerns (your prep pipeline, or prepare_documents.py when it lands) — ingest.py trusts what you hand it.
Invocation:
uv run --script scripts/ingest.py \
--data processed.jsonl \
--index <index_name> \
--sentinel-field <fts_field>
Flags:
| Flag | Short | Required | Purpose |
|---|---|---|---|
--data |
-d |
yes | Path to JSONL file with prepared documents (one per line) |
--index |
-i |
yes | Pinecone index name (must already exist) |
--sentinel-field |
-f |
yes | An FTS-enabled field on the index, used for the readiness-poll query. Pick the longest free-text field on your schema. |
--namespace |
-n |
no | Default __default__ |
--batch-size |
-b |
no | Default 50 (matches the SDK's own batch_upsert default). Reduce for large dense vectors. A 50-doc batch with 3072-dim float vectors lands ~5-10 MB and can be rejected; drop to --batch-size 25 (or lower) at high dimensions. |
--max-concurrency |
— | no | Default 4. Parallel HTTP connections used to upload batches. |
--poll-deadline |
— | no | Default 300 (seconds). Time to wait for documents to become searchable before giving up. |
--sentinel |
-s |
no | Token used for the readiness-poll query. Default: first whitespace-separated token of doc[0][sentinel-field]. |
What the script prints:
Loading processed.jsonl ...
Loaded 5000 document(s).
Sentinel: body='The'
Upserting in batches of 50 ...
batch @ 0: 50 docs in 0.31s (total: 50/5000)
batch @ 50: 50 docs in 0.29s (total: 100/5000)
...
Upsert complete: 5000 doc(s) in 21.4s.
Polling for searchability (deadline 300s) ...
Searchable after 12.3s (3 probe(s)).
Done — total 33.7s.
If a batch fails, the script prints every error message and exits non-zero. If the poll deadline expires, the script prints a hint about why (sentinel field isn't FTS-enabled, deadline too tight, docs structurally upserted but rejected by the inverted-index builder) and exits non-zero. Don't suppress these errors — they're surfacing real problems with the data or the index.
When you should NOT use the script:
- The user is doing per-doc patch updates. Use
documents.update(...)for partial field updates (see Updating documents inreferences/ingestion.md) — the script is for bulk loads, not per-record operations. - The user is ingesting from a non-JSONL source (CSV, Parquet, Postgres dump). Convert to JSONL first; the script doesn't parse other formats.
- The user explicitly asks you to write the ingestion code from scratch (teaching context). Honor the request and follow the canonical pattern:
documents.batch_upsert+result.has_errorsinspection +documents.searchpolling with sentinel and deadline.
The script lives at scripts/ingest.py relative to this skill directory. PEP 723 inline-metadata script — uv run --script installs typer and pinecone automatically on first invocation. No setup needed.
Use cases
Three concrete shapes to model your task on. Match the user's request to the closest one and follow its steps; improvise if the task is genuinely a hybrid.
UC-1: Index a new corpus end-to-end
Trigger. "Index this CSV / JSONL / folder for search," "build a search backend over [my articles / products / tickets / transcripts]," "make my [dataset] searchable."
For unprocessed / messy data, load the onboarding walkthrough first. If the user is showing up with raw data (unclear field types, possibly long text fields exceeding FTS limits, comma-separated tag strings, dates as strings, possibly duplicate IDs, etc.) and they haven't given you an explicit schema, read references/onboarding-walkthrough.md and follow it stage-by-stage. It's a conversational guide — meet the data, surface the processing decisions to the user, propose a schema, confirm before creating, then process+ingest+verify together. The walkthrough exists because schemas are immutable and "onboarding a new corpus" is a high-stakes flow that benefits from explicit user buy-in at each decision point.
If the user already gave you a clean JSONL + a schema spec, follow the abbreviated steps below.
Steps (when data is already prepared and the schema is decided):
- Inspect the corpus shape — text fields, structured metadata, do you also need a vector? Match it to one of the canonical shapes in
references/schema-design.md(articles, products, tickets, image library, code). - Pick analyzer settings on each text field —
language,stemming,stop_words. Stemming on for long prose, off for proper nouns / identifiers. Decide which fields are FTS and which are filterable-only metadata — see Filterable metadata isn't declared in the schema below; filterable-only metadata is a documents-side decision, not a schema field. - Assemble the schema with
SchemaBuilderand confirm it with the user before callingindexes.create— schemas are immutable in2026-07, so a wrong call costs a re-ingest. - Create the index.
pc.indexes.create(...)polls until the index is ready by default — no separate wait loop needed unless you passedtimeout=-1. - Run
scripts/ingest.py --data <jsonl> --index <name> --sentinel-field <fts_field>— see the Ingesting — use the packaged helper section above. The script handlesbatch_upsert+ per-batch error inspection + post-upsert readiness polling in one invocation. Don't hand-write the loop unless the user explicitly asks you to. - (The script polls automatically — by the time it exits cleanly, the index is searchable. If you skip the script and roll your own, you must poll
documents.searchwith a sentinel query and a deadline;batch_upsertreturning ≠ searchable — this is a document-indexing wait, separate from and in addition to the index-creation wait in step 4.) - Validate with one or two probe queries against fields you know contain the sentinel content.
Result. A working documents.search call against the user's data, returning ranked matches.
UC-2: Add a dense (or sparse) signal to a text-only corpus
Trigger. "Add semantic search," "add embeddings," "make this hybrid," or any prompt that describes a query pattern text alone can't serve (visual similarity, mood, cross-modal "looks like").
Steps.
- Confirm the new signal represents a modality or signal text can't express — image / audio / external score, or a different corpus than the existing FTS field. Re-encoding the same text into a dense field is an anti-pattern (
references/schema-design.md→ "When to add a dense field at all"). - Because schemas are immutable, plan a new index, not a migration. Get user confirmation before recreating. A hybrid index must declare its
sparse_vectorfield explicitly at creation — there is no way to add one later. - Pick an embedding provider and pin its output dimension at schema time. Beware payload-size pitfalls at native dimensions — Gemini-3072 etc. need truncation (
references/ingestion.md→ "Dense-vector payload size"). - Schema → create (blocks until ready by default) → ingest with embeddings inline or pre-cached.
- Validate with a hybrid query:
dense_vectorscore_by + text-match filter ($match_phrase/$match_all). That's the supported single-call cross-modal shape.
Result. One index, two retrieval shapes — pure text and dense+filter hybrid — both runnable without further setup.
UC-3: Build a documents.search call from a natural-language user prompt (agent mode)
Trigger. Agent receives a user prompt like "find articles about machine learning that mention TensorFlow and were published after 2024" or "documents about climate policy ranked by similarity to this paragraph." The index already exists.
Steps.
- (Optional) Discover the schema by calling
pc.indexes.describe(<NAME>)and readingschema.fields. Skip if you already know the field types from earlier in the conversation. - Decompose the user's prompt into
score_by/filtershapes using the agent-mode decomposition table below. (Hard requirements →filter. Ranking signals →score_by. Always includeinclude_fieldsexplicitly.) - Construct the
documents.search(...)call following the rules in the Querying section above — one scoring type per request, operator/field-type matching,include_fieldsalways set. - Execute the call. The response carries
resp.matches; iterate to getm._id,m._score, and field values viam.to_dict(). Use the matches in whatever shape the user asked for. - If results come back empty or wrong, walk the failure tree in
Common gotchas.
Result. Live search results matching the user's intent.
The four common UC-3 mistakes to actively avoid:
- Mixing scoring types in
score_by(server rejects). Put hard requirements infilter; rank by one signal inscore_by. - Putting hard requirements in
score_byas BM25 terms instead of infilteras$match_all/$match_phrase(returns ranked results that don't guarantee the term is present). - Operator/field-type mismatches (e.g.
$match_allon a float field,$gton a string field). Consult the operator table in the Querying section. - Omitting
include_fields(some SDK/backend builds 400). Always pass it explicitly.
Agent-mode query decomposition
Map user prompt cues to API shapes. Read top-down — identify the cue, copy the corresponding shape.
| User prompt cue | API shape |
|---|---|
| Open-ended keywords ("articles about machine learning", search-bar query) | score_by=[{"type": "text", "field": "<field>", "query": "<terms>"}] — BM25 token-OR |
| Exact phrase, drives ranking ("rank by 'beautifully written'") | score_by=[{"type": "query_string", "query": '<field>:("phrase here")'}] |
| Exact phrase, hard requirement ("must contain 'machine learning'") | filter={"<field>": {"$match_phrase": "machine learning"}} |
| Required tokens, any order ("must mention TensorFlow", "must be about Illinois") | filter={"<field>": {"$match_all": "tokens space-separated"}} — preferred over query_string +token because it's a true hard filter, doesn't contribute to score |
| At least one of these tokens ("contains AI or ML or robotics") | filter={"<field>": {"$match_any": "AI ML robotics"}} |
| Excluded tokens ("not about deprecated", "no opinion pieces") | filter={"$not": {"<field>": {"$match_any": "deprecated opinion"}}} — or -token inside query_string |
| Boolean / boost / slop / phrase-prefix ("weight 'eagle' 3x", "within N words") | score_by=[{"type": "query_string", "query": '<expr with ^N / ~N / "…"*>'}] — only Lucene supports these |
| Cross-field boolean ("title or body contains X") | score_by=[{"type": "query_string", "query": 'title:(X) OR body:(X)'}] |
| Numeric / date / range / boolean metadata ("after 2024", "rating > 4", "in stock") | filter={"<field>": {"$gt": ..., "$gte": ..., "$eq": ..., "$exists": true}} |
| Category / tag / list membership ("category = fiction", "tagged X") | filter={"<field>": {"$in": [...]}} (works on plain filterable metadata and string_list filterable fields) |
| Semantic similarity / mood / topic ("articles about ML", "documents that feel sombre") | score_by=[{"type": "dense_vector", "field": "<embedding_field>", "values": embed(<text>)}] — requires a dense_vector field |
| Visual appearance / cross-modal text query against an image corpus | Same dense_vector shape, with the embedding model that produced the stored image vectors. Multimodal embedders (Gemini-2 etc.) map a text query into the image space. |
| Hybrid: lexical requirement + semantic ranking ("articles about ML that mention TensorFlow") | Lexical → filter ($match_all / $match_phrase); semantic → score_by (dense_vector). Single call. |
Two structural rules the agent must enforce, no exceptions:
- One scoring type per request.
score_byacceptstext/query_string/dense_vector/sparse_vector, but a request ranks by one. Don't mix dense + text inscore_by— the server rejects it. Multi-field BM25 is the only "list" pattern that's allowed (multipletextclauses, or one cross-fieldquery_string). - Hybrid = filter + score_by, not two
score_byclauses. When a prompt has both a lexical requirement and a semantic ranking signal, lexical goes infilter(via$match_*operators) and semantic goes inscore_by. If both signals genuinely need to drive ranking, run two searches and merge IDs client-side.
Filterable metadata isn't declared in the schema at all
This is the single biggest shape change from the old pinecone.preview API, and it's easy to get only half right.
On a managed index (the deployment type every example in this skill uses — {"deployment_type": "managed", "cloud": ..., "region": ...}, which is also the default when deployment= is omitted), the schema may only declare fields that participate in search: dense_vector, sparse_vector, and string fields with full_text_search enabled. Every other field type — string (filterable, no FTS), string_list, float, boolean — is rejected at create time with a 400 if it appears in the schema. This is confirmed live, not just documented: the server's own error names all four types explicitly — "The schema only accepts fields used for search (field types dense_vector, sparse_vector, and string with full_text_search configuration). To use field '<name>' for filtering (field types boolean, float, string, or string_list), omit it from the schema and include it in documents." (That restriction is specific to managed/BYOC deployments — schema-declared filterable metadata is only legal on pod deployments, which this skill doesn't cover.) The SchemaBuilder methods add_float_field, add_boolean_field, and add_string_list_field still exist and still work correctly for a pod deployment; for the managed deployments this skill always uses, don't call any of them.
Instead: don't declare any filterable-only field in the schema, of any type. Just include the field in the documents you upsert — Pinecone indexes whatever's present on an upserted document for filtering automatically (exact-match on strings and numbers/booleans, membership on lists), whether or not it appears in the schema, with no configuration needed.
# WRONG on a managed index — the server 400s on EVERY one of these, not just category:
schema = (
SchemaBuilder()
.add_string_field("body", full_text_search={"language": "en"})
.add_string_field("category", filterable=True) # <-- rejected
.add_float_field("year", filterable=True) # <-- also rejected
.add_string_list_field("tags", filterable=True) # <-- also rejected
.build()
)
# RIGHT — the schema declares only search fields. category/year/tags are
# simply included on upserted documents and auto-index for filtering.
schema = (
SchemaBuilder()
.add_string_field("body", full_text_search={"language": "en"})
.build()
)
idx.documents.upsert(namespace=NAMESPACE, documents=[{
"_id": "doc-1",
"body": "...",
"year": 2025.0, # not in the schema — still filterable via $gt / $gte / $eq
"tags": ["classic"], # not in the schema — still filterable via $in / $nin
"category": "fiction", # not in the schema — still filterable via $eq / $in / $exists
}])
The forward-looking note in references/schema-design.md from the old preview docs — "declare metadata fields today to be future-proof" — no longer applies; declaring any filterable-only field is now actively wrong, not just unnecessary.
Workflow at a glance
Three phases. Each has its own reference file — consult it before writing code for that phase.
- Design the schema. Decide which string fields are full-text-searchable (declared in the schema), whether you need a
dense_vectorfield (and whether it earns its place), and whether you also need asparse_vectorfield — those are the only things that ever go in the schema. Every filterable metadata field — string, string_list, float, boolean alike — is NOT declared; it's just included on upserted documents. Schemas are fixed at index creation in2026-07— plan carefully. →references/schema-design.md - Ingest documents. For bulk loads from a prepared JSONL, run the bundled
scripts/ingest.pyhelper (it doesbatch_upsert+ error inspection + readiness polling correctly by construction — see the Ingesting — use the packaged helper section above). For per-doc patch updates, usedocuments.update(...). Either way, documents are indexed asynchronously after the HTTP call returns;batch_upsertreturning 202 ≠ searchable. →references/ingestion.mdfor the canonical pattern in detail. - Query the index. A single search request ranks by one scoring type — pass exactly one of
text,query_string,dense_vector, orsparse_vectorinscore_by(multi-field BM25 is supported via multipletextclauses or a cross-fieldquery_string). Layerfilter={...}for text-match ($match_phrase/$match_all/$match_any) and metadata filters ($eq/$in/$gte/$exists/$and/$or/$not). Control the response payload withinclude_fields. →references/querying.md
Quick template
End-to-end skeleton for a minimal text + filterable-metadata index. Copy it and edit every spot marked # TODO:. The template deliberately omits external embedding calls so it stays generic; see references/ingestion.md for dense / sparse field patterns and embedding-provider integration, and references/querying.md for the four scoring shapes plus text-match and metadata filters.
import time
from pinecone import Pinecone, SchemaBuilder
INDEX_NAME = "my-fts-index" # TODO: name your index (lowercase alphanumeric + hyphens, 1-45 chars)
NAMESPACE = "__default__" # TODO: pick a namespace; auto-created on first upsert
pc = Pinecone() # reads PINECONE_API_KEY
# TODO: preprod backends require an x-environment header on the client:
# pc = Pinecone(additional_headers={"x-environment": "preprod-aws-0"})
# 1. Schema — one FTS string field. That's the only kind of field that goes
# here: `category` and `year` below are deliberately NOT in the schema —
# see "Filterable metadata isn't declared in the schema at all" in
# SKILL.md. Field names must NOT start with `_` (reserved for `_id` /
# `_score`) or `$` (reserved for filter operators), and are limited to 64
# bytes.
schema = (
SchemaBuilder()
.add_string_field("body", full_text_search={"language": "en"}) # TODO: rename for your content
.build()
)
# 2. Create the index. Polls until ready by default (pass timeout=-1 to
# return immediately instead). Deployment defaults to managed/aws/us-east-1
# when omitted; pass `deployment=` explicitly to pick a different region.
# read_capacity defaults to {"mode": "OnDemand"}; pass
# {"mode": "Dedicated", ...} only if you specifically want provisioned reads.
if not pc.indexes.exists(INDEX_NAME):
pc.indexes.create(
name=INDEX_NAME,
schema=schema,
deployment={"deployment_type": "managed", "cloud": "aws", "region": "us-east-1"},
)
idx = pc.index(name=INDEX_NAME)
# 3. Upsert a single document. `_id` is required, every other field is optional.
# `category` and `year` aren't in the schema but are still filterable —
# see above.
# upsert REPLACES the document on conflict; use documents.update(...) for
# per-field patches (references/ingestion.md).
idx.documents.upsert(
namespace=NAMESPACE,
documents=[{
"_id": "doc-1",
"body": "Full-text search is great for keyword queries.",
"category": "intro",
"year": 2025.0,
}],
)
# 4. Poll until the FTS side is searchable (upsert returns BEFORE docs are indexed).
deadline = time.time() + 300
while time.time() < deadline:
resp = idx.documents.search(
namespace=NAMESPACE, top_k=1,
score_by=[{"type": "text", "field": "body", "query": "search"}], # TODO: sentinel query likely to hit
include_fields=[], # required on every search; [] = lightest payload (ids + _score only)
)
if resp.matches:
break
time.sleep(5)
# 5. Search — text scoring composed with metadata filter.
resp = idx.documents.search(
namespace=NAMESPACE,
top_k=5,
score_by=[{"type": "text", "field": "body", "query": "keyword queries"}],
filter={"year": {"$gte": 2024}}, # TODO: adjust filter or drop it
include_fields=["*"], # "*" = all stored fields; [] = `_id` + `_score` only
)
for m in resp.matches:
print(m._id, m._score, m.to_dict())
Common gotchas
- No filterable metadata field goes in the schema on managed indexes — string, string_list, float, and boolean alike. Only
dense_vector,sparse_vector, and FTS-enabledstringfields are legal inschema=; every other field type is rejected with400. Confirmed live against the real API, not just documented. Omit it and let it auto-index from upserted documents instead. See Filterable metadata isn't declared in the schema at all above — this is the change most likely to break code carried over from the oldpinecone.previewAPI, where such fields were declarable. - One scoring type per search request.
score_byacceptstext,query_string,dense_vector, orsparse_vector— but a request ranks by one type. Multi-field BM25 is fine (pass severaltextclauses, or a single cross-fieldquery_string). To combine BM25 ranking with adense_vector(orsparse_vector) signal, restrict the dense search with a text-matchfilteroperator ($match_phrase/$match_all/$match_any) on the lexical field, not by mixing types inscore_by. The "blend a dense vector and a text clause inscore_by" pattern is rejected by the server. - Text-match filter operators are the cross-modal hinge.
$match_phrase(exact phrase),$match_all(every token, any order),$match_any(at least one token) are filter-side operators onfull_text_searchfields. Each takes a single string (max 128 tokens). They reuse the field's tokenizer / stemmer, compose under$and/$or/$not, and are the supported way to compose lexical pre-filtering with dense or sparse ranking. Phrase slop ("…"~N), term boost (^N), and phrase prefix ("… word"*) are scoring-only — they live inquery_string, not infilter. - Preprod backends need
additional_headers={"x-environment": "..."}on thePinecone()client. Missing the header lands you on prod and you'll see "index not found" / empty-result symptoms that look like code bugs but aren't. include_fieldsis required on everydocuments.search(...)call. Pass["*"]for all stored fields or a list of names to project. Omitting it on some SDK/backend builds yields400instead of a sane default; always pass it explicitly to avoid surprises.- Match score is
_score; doc id is_id. The system match score is always on the_scorefield so a user metadata field literally namedscorecan coexist. Always readm._score, neverm.score. - Reserved field names: leading
_and$, max 64 bytes._is for system fields (_id,_score);$is for filter operators. Schema validation rejects names that violate either rule. Length cap is bytes, not characters — be careful with non-ASCII names. - Vector-field cardinality: at most one
dense_vectorand at most onesparse_vectorper index in2026-07. Multiple text fields are fine. - A hybrid index must declare its
sparse_vectorfield at create time — there's no adding one later.metric="dotproduct"on the dense field is NOT a hybrid declaration by itself in2026-07(it was in the old preview API). The create call succeeds either way; if thesparse_vectorfield is missing, only the sparse writes are refused, later, often from a different part of the codebase. If you're porting an old preview schema, audit everymetric="dotproduct"dense field for a missing sparse field before recreating it. batch_upsertfailures are silent by default. The return value carrieshas_errors,failed_batch_count, and a list ofBatchErrorobjects witherror_message. If you don't inspect them, you'll see "Uploaded 0 / N" and an indefinite "not yet indexed" poll — with the real cause (payload-too-large, schema mismatch, reserved field name) hidden. Always printresult.errors[*].error_messagebefore downstream steps.- Dense-vector payload size matters at batch time. A 50-doc batch with 3072-dim float vectors lands around 5–10 MB and can be rejected. If every batch fails, try reducing the embedding dimension via your provider's truncation knob (e.g. Gemini's
output_dimensionality=768) before debugging schema. - Async indexing:
batch_upsertreturning ≠ searchable. The server builds inverted indexes in the background after the HTTP call returns. If you query immediately you'll see empty result sets. Always polldocuments.searchwith a sentinel query and a deadline (pattern inreferences/ingestion.md). This is separate from — and in addition to —pc.indexes.create()'s own default polling for the index becoming ready. - String FTS field shape is
full_text_search={...}(dict), orTruefor server defaults. User-settable sub-fields:language,stemming,stop_words,ngram. Server-applied (visible indescribe()responses but NOT settable at index creation):lowercase(defaulttrue) andmax_term_len. Stemming is opt-in (defaultfalse) and required ifstop_words=Trueis set. A string field is either FTS-enabled or filterable, never both on a managed index — passingfilterable=Truealongsidefull_text_searchmakes the server silently keep the filter and drop the search config. - Schemas are fixed at index creation in
2026-07. Adding, removing, or retyping fields after creation is not supported. Changing dimension or metric on an existing vector field requires a new index. Plan the schema once. - Per-field updates are supported:
documents.update(...). Passdocuments=[{"_id": ..., "set_fields": {...}}]-style records, orfilter=+set_fields=/remove_fields=to patch many documents at once by metadata match.documents.upsertstill fully replaces a document on conflicting_idif that's what you want instead. Seereferences/ingestion.md→ "Updating documents". - Document operations: search, fetch, and delete all support
filternow.documents.fetchanddocuments.deleteboth gained afilterparameter — pass exactly one ofids,filter, or (for delete)delete_all. A filteredfetchis paginated (up to 10,000 docs per page viapagination_token); an ID-based fetch is never paginated.documents.deletereturns a response object withmatched_records(a point-in-time count for filtered deletes;Nonefor ID-list ordelete_alldeletes — the delete itself is applied asynchronously). documents.list(...)enumerates document IDs in a namespace, with no equivalent in the old preview API. Lazily-paginated, sorted by ID, optionally filtered byprefix. Seereferences/querying.md→ "documents.list— enumerate document IDs".- Namespaces auto-create on first upsert. Pass any namespace string to
documents.upsert/batch_upsertand the namespace is created on the fly; documents from different namespaces are fully isolated. Use"__default__"if you don't need partitioning. - Namespace management and
describe_index_statsnow work on document-schema indexes.idx.create_namespace(name=...),idx.list_namespaces(),idx.describe_namespace(name=...),idx.delete_namespace(name=...), andidx.describe_index_stats()are all available and confirmed working — seereferences/ingestion.md→ "Namespace management" for signatures and examples. - Document and request size limits: per-document max 2 MB; per-request max 2 MB and 1000 documents; per FTS-enabled
stringfield max 100 KB and 10,000 tokens (tokens > 256 bytes are truncated by the analyzer); per-document filterable metadata (everything not in an FTS field) max 40 KB. A schema can declare up to 100 FTS string fields. For long-prose corpora, chunk before ingest — seereferences/ingestion.md. score_byclause shape — singularfieldis canonical fortext/dense_vector/sparse_vector; onlyquery_stringtakes afieldsarray.text:{"type":"text", "field":"<fts_field>", "query":"<terms>"}.query_string:{"type":"query_string", "query":"<lucene>", "fields":["<a>","<b>"]}(the optionalfieldsarray;query_stringalso accepts a bare"fields":"body"string and the legacy"field":"body"as an alias).dense_vector:{"type":"dense_vector", "field":"<dense_field>", "values":[/*floats*/]}.sparse_vector:{"type":"sparse_vector", "field":"<sparse_field>", "sparse_values":{"indices":[...],"values":[...]}}— notesparse_values(NOTvalues) for sparse clauses.
- Single-term prefix wildcards aren't supported.
auto*doesn't work inquery_string; use phrase prefix ("machine lea"*— phrase must contain at least two terms, last term is matched as prefix). - Indexes can't be created in CMEK-enabled projects alongside any
full_text_searchfield, no backup/restore, no fuzzy or regex search, no S3 bulk import for document-shaped indexes in2026-07. If any of these are hard requirements, the document-schema FTS surface isn't yet ready.
Extension points
Currently shipped under scripts/:
scripts/ingest.py— bulk-ingest a prepared JSONL into an existing FTS index. Handlesbatch_upsertin safe-sized chunks, inspects every batch'sresult.errorsand aborts loudly on failure, then pollsdocuments.searchwith a sentinel + deadline until docs are searchable. Schema-agnostic: takes only--data,--index,--sentinel-field. Usage in Ingesting — use the packaged helper section above.
Query construction does NOT have a packaged helper — write documents.search(...) calls directly per the Querying section above.