Use built-in chat() to ask questions about your documents with your preferred LLM.
- Set up the PageIndex client — choose your index mode and LLM.
- Send a query — prepare messages and select documents.
- Streaming — stream the response as it is generated.
- Citations — include source references in the response.
Set up the PageIndex client
Two settings need to be configured:
index— determines where documents are processed and stored.chat— specifies the model that searches the tree and writes the answer.
Cloud
PageIndex API key, for indexing:
export PAGEINDEX_API_KEY="your-pageindex-key"LLM provider API key, for chat:
export OPENAI_API_KEY="your-openai-key"from pageindex import PageIndexClient
client = PageIndexClient(
index="cloud", # index and store in PageIndex Cloud
chat="gpt-5.6-sol", # model that searches the tree
)Customize system prompt
Pass instructions= when creating the client to customize the answering agent across questions. These instructions are appended after PageIndex’s default system prompt.
client = PageIndexClient(
index="cloud",
chat="gpt-5.6-sol",
instructions="You are ACME's support assistant. Answer in the user's language.",
)Use an existing document ID from client.list_documents(), or submit a new document.
doc_id = client.submit_document("./2023-annual-report.pdf", wait=True)["doc_id"]The examples above use OpenAI. Any provider works — see Use different LLMs.
Send a query
Start with a question string or a list of conversation messages.
Single message
messages = "What are the key findings in this document?"Then choose the documents or folder and send messages:
Document
answer = client.chat(messages, doc_id="doc_id_1")Streaming
Streamed answers show the agent’s process by default — thinking and tool calls woven into the text as it runs:
for chunk in client.chat("Summarize this document", doc_id=doc_id, stream=True):
print(chunk, end="", flush=True)[thinking] I should read the report's key sections first.
[tool_call] get_document_structure {"doc_name": "report.pdf"}
[tool_result] get_document_structure: {"success": true, ... (+304 chars)
The report finds that ...Use show_process=False to stream only the answer.
client.chat("...", doc_id=doc_id, stream=True, show_process={"thinking": False})To customize the process output, pass a dict to show_process:
| Key | Description | Default |
|---|---|---|
thinking | Show the model’s thinking when available. | True |
tool_call | Show tool calls. | True |
tool_result | Show tool results. | True |
max_chars | Character limit per process summary line. | 200 |
Citations
Pass citations=True to request source citations. Citations use <cite doc="…" page="…"/> tags, with block="…" where the cloud document has block-level data:
answer = client.chat("Summarize the document.", doc_id=doc_id, citations=True)Revenue increased during the reporting period. <cite doc="report.pdf" page="12"/>Block-level citations — cloud only. Local citations stop at the page. On cloud, a citation also carries a block attribute naming one layout block:
Revenue increased during the reporting period. <cite doc="report.pdf" page="12" block="p12_text_3"/>Read the citations
get_citations() parses the tags in an answer into a list, resolving each doc name to a document ID and each block to its position:
citations = client.get_citations(answer)
# [{'document': 'report.pdf', 'doc_id': 'pi-…', 'page': 12,
# 'block_id': 'p12_text_3', 'bbox': [78, 24, 293, 44], 'block_type': 'text', 'text': '…'}]block_id appears on block-level citations; bbox, block_type, and text come with it when the block could be read — a local document or an unreadable block has none, so use citation.get("bbox") rather than indexing.
bbox is [x0, y0, x1, y1] in thousandths of the page’s width and height, origin top-left.
Display the citations
resolve_citations() returns the same data ready to render: the answer with each tag replaced by a numbered markdown link, and one citation entry per distinct source, each carrying its anchor and index.
resolved = client.resolve_citations(answer)
print(resolved["answer"])
# Revenue increased during the reporting period. [[1]](#pageindex-citation-01)
for c in resolved["citations"]:
print(c["index"], c["anchor"], c["document"], c["page"])
# 1 pageindex-citation-01 report.pdf 12The same source cited twice reuses its number, and your page renders the anchor targets from the anchor field.
Highlight the cited region
Fetch the cited page as an image and draw the block’s box on it:
import requests
from pageindex import highlight_region
c = resolved["citations"][0]
url = client.get_page_image(c["doc_id"], c["page"])
image = highlight_region(requests.get(url).content, c["bbox"])
image.save("highlighted_page.png")get_page_image() returns a short-lived URL to the page as a JPEG, and get_document_image(doc_id, "img-7.jpeg") does the same for an image OCR extracted from the document. Both are cloud-only.
API protocols
Use protocol= to return a protocol’s native response format. With stream=True, the response is its native event stream.
Chat Completions
response = client.chat(
"Summarize this document",
doc_id=doc_id,
protocol="chat_completions",
)
print(response["choices"])Returns the Chat Completions envelope with choices and usage. This protocol also works with managed cloud chat, without your own chat model. chat_completions() remains available for existing code.
show_process does not apply when protocol is set. Pass provider-specific fields such as thinking and top_k through extra_body.
chat() parameters
chat() parameters| Name | Type | Required | Description | Default |
|---|---|---|---|---|
| messages | string or List[Dict] | yes | A question string, or role/content conversation history. | - |
| doc_id | string or List[string] | no | Document ID(s) to scope the conversation. | None |
| folder_id | string | no | Cloud-only — steer discovery toward this folder’s documents; set to "root" or omit both folder_id and doc_id to search the whole library. | None |
| stream | boolean | no | Yield the answer as text chunks as it is produced; with a protocol, that protocol’s native event stream. | False |
| show_process | boolean or Dict | no | Streamed chat: weave the run into the text. False for the bare answer; dict keys thinking / tool_call / tool_result / max_chars. | on |
| model | string | no | Backend model name; defaults to the client’s chat_model (with protocol="messages", only one you set yourself). | None |
| reasoning_effort | string | no | "low" / "medium" / "high", sent in each lane’s native spelling. | None |
| protocol | string | no | "chat_completions", "responses" or "messages": drive that wire natively, with its own input/output shapes. The managed cloud chat serves "chat_completions" too. | None |
| instructions | string or List[Dict] | no | Guidance for this call, appended after the managed system prompt and the client’s own instructions (Messages system blocks with protocol="messages"). | None |
| citations | boolean | no | Request source citations with <cite doc= page=/> tags (block= where the cloud document has blocks); on the managed cloud chat, the endpoint’s own citations. | False |
| max_turns | int | no | Cap on agent turns per call. | None |
| backend | Dict | no | Per-call connection overrides, merged over the client’s chat_backend. | None |
| extra_headers | Dict | no | Extra HTTP headers on each backend request. | None |
| extra_body | Dict | no | The provider’s own request fields beyond these parameters, merged last. | None |
model, reasoning_effort, max_turns, backend, extra_headers and the responses / messages protocols apply to own-model chat (a configured chat_model); the managed cloud chat selects its own model and rejects them. instructions, citations, extra_body and protocol="chat_completions" work on both.