Skip to Content
Introducing PageIndex Flash

PageIndex LLM Integration

Use built-in chat() to ask questions about your documents with your preferred LLM.

  1. Set up the PageIndex client — choose your index mode and LLM.
  2. Send a query — prepare messages and select documents.
  3. Streaming — stream the response as it is generated.
  4. Citations — include source references in the response.

Set up the PageIndex client

Two settings need to be configured:

  • index — determines where documents are processed and stored.
  • chat — specifies the model that searches the tree and writes the answer.

PageIndex API key, for indexing:

export PAGEINDEX_API_KEY="your-pageindex-key"

LLM provider API key, for chat:

export OPENAI_API_KEY="your-openai-key"
from pageindex import PageIndexClient client = PageIndexClient( index="cloud", # index and store in PageIndex Cloud chat="gpt-5.6-sol", # model that searches the tree )

Customize system prompt

Pass instructions= when creating the client to customize the answering agent across questions. These instructions are appended after PageIndex’s default system prompt.

client = PageIndexClient( index="cloud", chat="gpt-5.6-sol", instructions="You are ACME's support assistant. Answer in the user's language.", )

Use an existing document ID from client.list_documents(), or submit a new document.

doc_id = client.submit_document("./2023-annual-report.pdf", wait=True)["doc_id"]

The examples above use OpenAI. Any provider works — see Use different LLMs.


Send a query

Start with a question string or a list of conversation messages.

messages = "What are the key findings in this document?"

Then choose the documents or folder and send messages:

answer = client.chat(messages, doc_id="doc_id_1")

Streaming

Streamed answers show the agent’s process by default — thinking and tool calls woven into the text as it runs:

for chunk in client.chat("Summarize this document", doc_id=doc_id, stream=True): print(chunk, end="", flush=True)
[thinking] I should read the report's key sections first. [tool_call] get_document_structure {"doc_name": "report.pdf"} [tool_result] get_document_structure: {"success": true, ... (+304 chars) The report finds that ...

Use show_process=False to stream only the answer.

client.chat("...", doc_id=doc_id, stream=True, show_process={"thinking": False})

To customize the process output, pass a dict to show_process:

KeyDescriptionDefault
thinkingShow the model’s thinking when available.True
tool_callShow tool calls.True
tool_resultShow tool results.True
max_charsCharacter limit per process summary line.200

Citations

Pass citations=True to request source citations. Citations use <cite doc="…" page="…"/> tags, with block="…" where the cloud document has block-level data:

answer = client.chat("Summarize the document.", doc_id=doc_id, citations=True)
Revenue increased during the reporting period. <cite doc="report.pdf" page="12"/>

Block-level citations — cloud only. Local citations stop at the page. On cloud, a citation also carries a block attribute naming one layout block:

Revenue increased during the reporting period. <cite doc="report.pdf" page="12" block="p12_text_3"/>

Read the citations

get_citations() parses the tags in an answer into a list, resolving each doc name to a document ID and each block to its position:

citations = client.get_citations(answer) # [{'document': 'report.pdf', 'doc_id': 'pi-…', 'page': 12, # 'block_id': 'p12_text_3', 'bbox': [78, 24, 293, 44], 'block_type': 'text', 'text': '…'}]

block_id appears on block-level citations; bbox, block_type, and text come with it when the block could be read — a local document or an unreadable block has none, so use citation.get("bbox") rather than indexing.

bbox is [x0, y0, x1, y1] in thousandths of the page’s width and height, origin top-left.

Display the citations

resolve_citations() returns the same data ready to render: the answer with each tag replaced by a numbered markdown link, and one citation entry per distinct source, each carrying its anchor and index.

resolved = client.resolve_citations(answer) print(resolved["answer"]) # Revenue increased during the reporting period. [[1]](#pageindex-citation-01) for c in resolved["citations"]: print(c["index"], c["anchor"], c["document"], c["page"]) # 1 pageindex-citation-01 report.pdf 12

The same source cited twice reuses its number, and your page renders the anchor targets from the anchor field.

Highlight the cited region

Fetch the cited page as an image and draw the block’s box on it:

import requests from pageindex import highlight_region c = resolved["citations"][0] url = client.get_page_image(c["doc_id"], c["page"]) image = highlight_region(requests.get(url).content, c["bbox"]) image.save("highlighted_page.png")

get_page_image() returns a short-lived URL to the page as a JPEG, and get_document_image(doc_id, "img-7.jpeg") does the same for an image OCR extracted from the document. Both are cloud-only.


API protocols

Use protocol= to return a protocol’s native response format. With stream=True, the response is its native event stream.

response = client.chat( "Summarize this document", doc_id=doc_id, protocol="chat_completions", ) print(response["choices"])

Returns the Chat Completions envelope with choices and usage. This protocol also works with managed cloud chat, without your own chat model. chat_completions() remains available for existing code.

show_process does not apply when protocol is set. Pass provider-specific fields such as thinking and top_k through extra_body.


chat() parameters

NameTypeRequiredDescriptionDefault
messagesstring or List[Dict]yesA question string, or role/content conversation history.-
doc_idstring or List[string]noDocument ID(s) to scope the conversation.None
folder_idstringnoCloud-only — steer discovery toward this folder’s documents; set to "root" or omit both folder_id and doc_id to search the whole library.None
streambooleannoYield the answer as text chunks as it is produced; with a protocol, that protocol’s native event stream.False
show_processboolean or DictnoStreamed chat: weave the run into the text. False for the bare answer; dict keys thinking / tool_call / tool_result / max_chars.on
modelstringnoBackend model name; defaults to the client’s chat_model (with protocol="messages", only one you set yourself).None
reasoning_effortstringno"low" / "medium" / "high", sent in each lane’s native spelling.None
protocolstringno"chat_completions", "responses" or "messages": drive that wire natively, with its own input/output shapes. The managed cloud chat serves "chat_completions" too.None
instructionsstring or List[Dict]noGuidance for this call, appended after the managed system prompt and the client’s own instructions (Messages system blocks with protocol="messages").None
citationsbooleannoRequest source citations with <cite doc= page=/> tags (block= where the cloud document has blocks); on the managed cloud chat, the endpoint’s own citations.False
max_turnsintnoCap on agent turns per call.None
backendDictnoPer-call connection overrides, merged over the client’s chat_backend.None
extra_headersDictnoExtra HTTP headers on each backend request.None
extra_bodyDictnoThe provider’s own request fields beyond these parameters, merged last.None

model, reasoning_effort, max_turns, backend, extra_headers and the responses / messages protocols apply to own-model chat (a configured chat_model); the managed cloud chat selects its own model and rejects them. instructions, citations, extra_body and protocol="chat_completions" work on both.

Last updated on