Understand Your AI Token Cost Before You Paste Again
Reduce repeated context, keep useful evidence, and estimate your AI token cost with assumptions you can change. Measure the difference in your own workflow.
What does the same policy question cost with less input context?
In this fictional API example, 20 pages at 500 tokens each supply 10,000 input tokens. At $2 per million input tokens, that context costs $0.02. Supplying a relevant 2,500-token excerpt costs $0.005 before outputs or retrieval. Those assumptions illustrate the arithmetic; they are not a measured Brain saving.
Policy context worksheet
Source excerpt · Policy context worksheet
Example only: 20 pages × 500 tokens = 10,000. Focused excerpt = 2,500. Input rate = $2 per million.
AI token cost grows when background travels with every task
A useful question can arrive with a surprisingly large attachment. You may need one clause from a policy, yet paste the entire handbook because finding the right paragraph takes longer. Repeat that across five conversations and ten teammates, and the same background consumes input tokens again. The context is familiar to you; the model still has to receive it for that request.
Start by separating three questions: how much text did you send, what does the selected model charge for it, and which allowance does your account use? An API bill is calculated differently from a consumer subscription’s usage window. A shorter input can change token usage, but it does not rewrite the provider’s subscription rules, reset a limit or remove the cost of the answer.
Brain makes a maintained knowledge collection available to compatible AI clients. Instead of reconstructing a document bundle for each conversation, ask a specific question and inspect the relevant sources. The goal is enough context to answer well. A tiny excerpt that omits an exception can be worse than a larger one, so cost experiments should compare answer quality alongside the token count.
Use the calculator below to establish a baseline for input context. Enter the pages you normally supply, the people repeating the workflow and a realistic focused-context percentage. Start a small pilot with a repeated task, record actual usage and check cited evidence. Keep retrieval, indexing, model outputs and Brain charges in the full evaluation even though this calculator intentionally separates them from pasted-context costs.
Connect. Ask. Govern.
From scattered documents to a shared answer
01
Connect the approved cost-test references
Choose the handbook passage and assumptions record for a representative repeated task. Exclude negotiated account terms or confidential records from the initial experiment, and record the source version used for the comparison.
02
Ask and inspect both answer quality and context
Compare a complete-document baseline with a focused source-backed answer. Check that the necessary exception remains present, then use actual token counts where available instead of treating pages as an exact tokenizer measurement.
03
Govern the experiment and its assumptions
Give test identities appropriate source scope, keep provider keys in approved settings and record which charges are included. Revisit the assumptions when the model, documents or retrieval behavior changes.
See the idea in action
ai token cost: questions with evidence
Fictional examples. These interactions do not query your Brain or test real permissions.
In this fictional API example, 20 pages at 500 tokens each supply 10,000 input tokens. At $2 per million input tokens, that context costs $0.02. Supplying a relevant 2,500-token excerpt costs $0.005 before outputs or retrieval. Those assumptions illustrate the arithmetic; they are not a measured Brain saving.
Policy context worksheet
Example only: 20 pages × 500 tokens = 10,000. Focused excerpt = 2,500. Input rate = $2 per million.
The fictional handbook asks travelers to attach an itemized receipt when claiming a meal expense. An answer can cite that section and its exception instead of sending every benefit, leave and payroll chapter. Check whether your question needs another section before deciding the excerpt is sufficient.
Travel receipt guidance
Meal claims require an itemized receipt. If a receipt is unavailable, attach an explanation for the reviewer.
The fictional project assumptions document records the pilot audience, reporting period and exclusions. A compatible client can consult that maintained source for the next task. This example does not increase a provider usage allowance or claim that all old conversation content is automatically saved.
Pilot assumptions record
Pilot scope: ten internal reviewers, one reporting month; external customer records excluded.
This fictional cost review contains negotiated commercial terms in a restricted appendix. Under limited access, the example withholds the appendix rather than offering it as context. A cost comparison should use approved test material and evaluate source access separately from the amount of text sent.
Restricted model agreement
Private appendix: negotiated account terms and billing contact details.
Open a question, then inspect its source excerpts.
Compare context costs in the client you already use
Use Claude, Cursor or Codex over MCP to consult a maintained policy or project note. Keep the task and model consistent when comparing context strategies, and check the client’s setup guide. Brain supplies knowledge; it does not change the provider’s usage allowance or make every model bill the same way.
Start the cost comparison with owned source material
Connect the Drive policies and Notion project notes your pilot is allowed to use, or upload a small document set. Keep the source version and task fixed so the comparison measures the context strategy. Drive and Notion are the available source connections; upcoming connectors are outside this pilot.
Google Drive policies
Use the handbook or policy folder behind the repeated question, retaining its source permissions.
Notion project notes
Consult maintained assumptions and decision pages rather than copying an entire project history into each prompt.
Uploaded pilot documents
Choose approved test files and record the page-to-token assumptions used in the worksheet.
Keep access intentional
Run the cost experiment on approved material
Cost testing should preserve the source boundary. Use synthetic or approved documents and keep restricted commercial terms out of an ordinary teammate’s context. The fictional comparison illustrates the input arithmetic; the actual access, provider processing and use of a model credential must be reviewed for the configured pilot.
Keep inherited permissions
A cheaper input is useful only if it is appropriate for the requester. Brain inherits access from connected sources: a policy appendix someone cannot open in Drive stays outside their answer.
Identify and scope the pilot agent
Each request is tied to a named person or agent. An agent key has its own scope, independently of its owner’s access, so a cost pilot can use a focused collection without opening every commercial record.
Review access without logging the text
The content-blind audit records who asked, what was reached or withheld, and when. Review access metadata alongside token measurements; the audit record does not contain the policy passages themselves.
Check model processing terms
Brain does not use connected material to train its own models. Managed answering sends permitted excerpts under provider terms that prohibit training; a model key you supply uses your agreement with that provider. Check the privacy policy before choosing a pilot configuration.
Fictional content-blind access record
Actor
Example cost-pilot agent
Source
Google Drive · fictional policy collection
Object
demo-policy-object-01
Timestamp
2026-09-01 09:00 UTC · example
Policy
Read within the pilot collection; restricted appendix excluded
Action
Read approved source
Outcome
Allowed within the example workspace
Fictional metadata only. No document content is shown, and this is not a record from your account or proof of live enforcement.
Choose a model input rate and compare the context supplied for the same number of tasks. The calculator separates uncached input context from outputs, retrieval, indexing, caching, Brain charges and taxes. Its estimates are a baseline to measure, not a fixed percentage reduction or a larger provider subscription allowance.
Estimate monthly input-context costs
Compare supplying full documents with a smaller relevant context. Choose the assumptions for your workload; this estimates uncached input tokens, not your total AI bill.
Full document context / month$22.0011,000,000 input tokens
Formula: sessions × working days × people × pages × tokens per page. Focused tokens use the percentage you choose. Costs multiply tokens by the selected price per million. Excludes output tokens, caching, retrieval and indexing, Brain charges, taxes and negotiated rates. Results are illustrative; actual context size and savings vary.
Provider prices checked 1 October 2026. Pricing source. These examples describe pricing, not the models available in your Brain account. See Brain plans and limits.
A straightforward baseline when the document is short or the task needs the whole text. Repeated large attachments can add input tokens and make it harder to notice which section supports the answer.
Use assistant project files
A useful way to organize work inside a provider’s project features. Providers can retrieve or manage context themselves; compare the actual behavior and usage on your account rather than assuming every file is resent in full.
Keep a long system prompt
Helpful for concise, stable instructions. A large prompt containing changing company facts is harder to maintain and may add repeated input. Separate durable instructions from documents that need source ownership and updates.
Consult Brain for relevant knowledge
Connect approved sources and inspect the passages behind a specific answer through compatible clients. Include retrieval, indexing, model output and Brain charges in the total comparison. Lower input context is a hypothesis to measure, not a universal percentage.
ai token cost: frequently asked questions
What is Claude's usage limit and why do I hit it?
Claude distinguishes usage limits across conversations from the length limit of an individual conversation. Your allowance depends on your plan and workload, including message length, attachments, model and features. Check your account’s current usage information. Brain does not change that allowance; focused context may reduce repeated input in a suitable task.
How are AI tokens counted and priced?
Tokens are units a model uses to process text and other supported inputs. Counts vary by tokenizer and content, so a page-to-token ratio is an estimate. API rates commonly separate input and output tokens, with caching or other charges depending on the provider. A subscription’s usage allowance is a different billing model.
How does Brain reduce token usage?
Brain lets compatible clients consult connected knowledge for a particular question instead of relying on a repeatedly pasted document archive. Relevant context can be smaller than the full collection. The result depends on the task and retrieval behavior. Measure actual tokens and include indexing, retrieval and answer generation when evaluating overall costs.
Does sending less context lower answer quality?
It can if you omit a relevant exception, definition or supporting section. Choose the context that answers the question accurately, then inspect the source evidence. Compare a focused answer with a complete-document baseline on representative tasks. A smaller token count is useful only when the answer remains adequate for the work.
Does Brain work with my Claude or ChatGPT subscription?
Brain connects through compatible MCP clients; current setup information names Claude, Cursor and Codex. Check the instructions for the client and account you intend to use. A provider subscription does not automatically cover every API or integration charge, and Brain does not increase its usage limits or guarantee ChatGPT account compatibility.
Can I use my own API key?
Published Brain setup information allows you to supply a model-provider key. Check the model configuration available in your account and the provider agreement that applies to it. Your key’s API usage can have separate billing and limits from a chat subscription. Never place an API key in a marketing form or shared document.
How do I estimate my own savings?
Enter your sessions, typical pages, team size and working days in the calculator. Set tokens per page and a focused-context percentage that reflects a plausible task. The result estimates uncached input costs using the chosen model price. Validate it against actual usage, then add output, retrieval, indexing and Brain charges for a complete budget.
Start with one repeated task
Compare full and focused context on an approved document, inspect the answers, and use real usage records to decide what belongs in your next workflow.