Skip to main content

AI & assistant-friendly summary

This section provides structured content for AI assistants and search engines. You can cite or summarize it when referencing this page.

Summary

A 5,000-document corpus on the site calculator is $163.28/month on S3 Vectors and $510.52/month on OpenSearch Serverless Classic. $160 of the smaller bill is Claude Sonnet 5.5.

Key Facts

  • •A 5,000-document corpus on the site calculator is $163.28/month on S3 Vectors and $510.52/month on OpenSearch Serverless Classic
  • •$160 of the smaller bill is Claude Sonnet 5.5
  • •On 8 October 2026, AWS Cost Explorer, AWS Budgets, and AWS Cost Management Dashboards added Amazon Bedrock product attributes
  • •You can group and filter Bedrock spend by model, model provider, inference type, and feature, at no extra charge, in every commercial Region
  • •Attribute history starts on 1 September 2026

Entity Definitions

Amazon Bedrock
Amazon Bedrock is an AWS service discussed in this article.
Bedrock
Bedrock is an AWS service discussed in this article.
S3
S3 is an AWS service discussed in this article.
Aurora
Aurora is an AWS service discussed in this article.
IAM
IAM is an AWS service discussed in this article.
OpenSearch
OpenSearch is an AWS service discussed in this article.
RAG
RAG is a cloud computing concept discussed in this article.
serverless
serverless is a cloud computing concept discussed in this article.

Bedrock Knowledge Base Cost Starts at the Vector Store

Quick summary: A 5,000-document corpus on the site calculator is $163.28/month on S3 Vectors and $510.52/month on OpenSearch Serverless Classic. $160 of the smaller bill is Claude Sonnet 5.5.

Key Takeaways

  • A 5,000-document corpus on the site calculator is $163.28/month on S3 Vectors and $510.52/month on OpenSearch Serverless Classic
  • $160 of the smaller bill is Claude Sonnet 5.5
  • On 8 October 2026, AWS Cost Explorer, AWS Budgets, and AWS Cost Management Dashboards added Amazon Bedrock product attributes
  • You can group and filter Bedrock spend by model, model provider, inference type, and feature, at no extra charge, in every commercial Region
  • Attribute history starts on 1 September 2026
Navy block with horizontal amber bands on a dark navy background, suggesting a bill split into layers
Table of Contents

On 8 October 2026, AWS Cost Explorer, AWS Budgets, and AWS Cost Management Dashboards added Amazon Bedrock product attributes. You can group and filter Bedrock spend by model, model provider, inference type, and feature, at no extra charge, in every commercial Region. GovCloud and the China Regions are excluded. Attribute history starts on 1 September 2026. A query whose start date is earlier fails with DataUnavailableException.

Those attributes show Bedrock lines such as on-demand inference and the reranker. They do not show OpenSearch Serverless or S3 Vectors. On a self-managed knowledge base, the omitted line is often the floor.

The Knowledge Base cost calculator uses us-east-1 rates dated 29 September 2026. One corpus: 5,000 documents, 16 KB average, weekly sync, 50,000 Retrieve calls, 10,000 RetrieveAndGenerate calls on Claude Sonnet 5.5 at 4,000 input tokens and 800 output tokens, rerank off, Data Automation off. S3 Vectors totals $163.28 a month. OpenSearch Serverless Classic totals $510.52. $160 of the smaller total is the model. The unrounded gap is $347.23.

Two prices, two products

The Bedrock pricing page, checked 9 October 2026, lists two Knowledge Base products.

Managed Knowledge Bases charge $5 per GB of raw data per month for the index. Standard Retrieve is $1 per 1,000 API calls. The managed parser, the managed embedding model, and the managed reranker are $0. Agentic retrieval adds $4 per 1,000 Agentic Retrieve calls plus $1 per 1,000 underlying Retrieve calls. If you pick your own embedding or rerank model, that model’s price is extra. The site calculator does not model this product.

Self-managed Knowledge Bases are the path the calculator does model. You choose the vector store. There is no separate hourly knowledge base fee. You pay embeddings, store, optional parsing, optional rerank, and the generation model.

If the 16 KB figure on this corpus is the raw size, 5,000 documents are about 0.08 GB. Managed index storage at $5 per GB is under $1. The 50,000 Retrieve calls are $50. Generation is still extra if you call a model afterward. Self-managed S3 Vectors, on the calculator’s illustrative query rate, prices 60,000 queries at $3. Those two query meters are not the same product. Pick managed when you want AWS to own parsing, embeddings, and the index, and the raw corpus stays small. Pick self-managed S3 Vectors when the documents already live in S3 and Retrieve volume would make $1 per 1,000 calls the line you notice.

Five meters on a self-managed knowledge base

MeterWhat the calculator usesThis corpus, weekly
EmbeddingsTitan Text Embeddings v2 at $0.02 per 1M tokens, 500 tokens per chunk$0.11
ParsingBedrock Data Automation standard output at $0.010 per page, off by default$0
Vector storeS3 Vectors illustrative $1/GB-month and $0.05 per 1,000 queries, or OpenSearch Classic$3.17 or $350.40
RerankAmazon Rerank 1.0 at $1 per 1,000 queries, off by default$0
GenerationClaude Sonnet 5.5 at $2 / $10 per 1M tokens$160.00

The 16 KB documents become 9 chunks each under the calculator’s rule (500 tokens, 4 bytes per token), so 5,000 documents are 45,000 vectors and 0.17 GB. Weekly sync re-embeds 25% of those chunks. Daily is 30%. A monthly cycle re-embeds 100%, which is $0.45 of Titan tokens. Moving weekly to daily changes embeddings from $0.11 to $0.14.

The $0.02 per million tokens, the S3 Vectors unit prices, and the $1 rerank rate are the calculator’s model as of 29 September 2026. The S3 Vectors storage and query rates are marked illustrative in that file. The $0.010 per page figure is the Bedrock pricing page’s own example for Data Automation standard output used as a Knowledge Base parser. Cohere Rerank 3.5, on that same page, is $2 per 1,000 queries, and one query holds at most 100 chunks. A request with 350 chunks is billed as 4 queries.

Sonnet 5.5 at $2 / $10 per million tokens is the rate in the site Bedrock calculator. That module treats it as Anthropic’s list price until the Bedrock pricing page prints the Bedrock rate. Re-check the model card before you lock a budget to $160.

OpenSearch compute is 2 OCUs × 730 hours × $0.24 = $350.40. The $0.24 rate is the site OpenSearch calculator. The OpenSearch pricing page, checked the same day, states the Classic minimum in words: at least 2 OCUs for the first collection (1 indexing OCU with primary and standby, 1 search OCU with a replica). The static page text states that minimum and leaves the hourly dollar in the regional price list. A few cents of managed storage on 0.17 GB sit inside the $510.52 total, so the rounded lines ($0.11 + $350.40 + $160.00) land one cent under it. That is also why $510.52 minus $163.28 is $347.24 while the unrounded gap is $347.23.

NextGen collections on that page have no minimum. Indexing and search OCUs scale to zero after 10 minutes of inactivity. A vector search collection still cannot share OCUs with search or time series collections, even on the same KMS key.

Read the Bedrock lines in Cost Explorer

Group a month of Bedrock by feature, then by model. The group type in the Cost Explorer API is PRODUCT_ATTRIBUTE. The keys documented for Bedrock are provider, model, inferenceType, and feature. Examples in the API reference include Claude Sonnet 5, Claude Haiku 4.5, Anthropic, input tokens, output tokens, on-demand inference, and reranker.

GetCostAndUsage does not require a service filter for this. Bedrock cost can show up under more than one service name. Omit the service filter so you do not drop a line. When you do filter, the service name has to match exactly, or the call fails validation.

AWS CLI 2 with Cost Explorer support for PRODUCT_ATTRIBUTE. The time period has to start on or after 1 September 2026. An older CLI that rejects the group type belongs in the console, which gained the same control on 8 October 2026.

aws ce get-cost-and-usage \
  --time-period Start=2026-10-01,End=2026-11-01 \
  --granularity MONTHLY \
  --metrics UnblendedCost \
  --group-by Type=PRODUCT_ATTRIBUTE,Key=feature

Filter one model the way the API reference does. Swap the model string for the name Cost Explorer shows, not the model ID.

aws ce get-cost-and-usage \
  --time-period Start=2026-10-01,End=2026-11-01 \
  --granularity MONTHLY \
  --metrics UnblendedCost \
  --filter '{"ProductAttributes":{"Key":"model","Values":["Claude Sonnet 5"],"MatchOptions":["EQUALS"]}}'

Budgets can filter on the same attributes. Build the budget in the console against one model, then attach it to the channel the platform team already reads. Combine the attributes with cost allocation tags on application inference profiles, and with IAM principal tags, when you need the application or the role and not only the model.

Then open the OpenSearch and S3 lines for the same month. If Bedrock attributes sum to roughly the generation total and OpenSearch is sitting near $350, the knowledge base is on a Classic collection.

Pick the store

Use S3 Vectors for a new self-managed knowledge base. The trade-off is hybrid keyword-plus-vector search, and the latency target already spelled out in the S3 Vectors post and the vector store decision. Those jobs stay on OpenSearch or MemoryDB. This page does not restate the June 2026 query limits.

If the job needs OpenSearch, create a NextGen collection and set a maximum OCU. Classic is the type that bills the 2 OCU floor the calculator models at $350.40 a month, including hours with no queries. Vector collections do not share that floor with a search collection on the same key. The first vector collection opens a new pair of OCUs.

Aurora PostgreSQL with pgvector is the right store when the application already runs on Aurora and the vectors should sit next to the rows. The calculator does not price it, so this page has no Aurora total to compare with the $163.28.

Sync the source bucket on change, not on a calendar that re-embeds a stable corpus. The calculator’s weekly assumption is 25% of chunks. A change of chunking strategy or embedding model rebuilds the index. On this corpus that rebuild is $0.45 of Titan tokens. Schedule it when the index definition changes.

Turn on Data Automation only for pages with tables, figures, or scans. The pricing page’s 1,000-page example at $0.010 is $10. Parsing all 5,000 documents on the weekly 25% assumption adds $12.50. Plain text does not need that parser.

Code that changes the bill

RetrieveAndGenerate always calls the model. Retrieve returns chunks and lets you skip generation when the question already has an answer.

boto3 bedrock-agent-runtime, Retrieve API. numberOfResults caps the chunks. The metadata filter runs against the index before that cap. Five chunks is the starting point for a chat answer.

client = boto3.client("bedrock-agent-runtime", region_name="us-east-1")

retrieved = client.retrieve(
    knowledgeBaseId=kb_id,
    retrievalQuery={"text": question},
    retrievalConfiguration={
        "vectorSearchConfiguration": {
            "numberOfResults": 5,
            "filter": {"equals": {"key": "department", "value": "engineering"}},
        }
    },
)

Attach rerank on the second call, after the first pass misses. Pull rerank_model_arn from the model card in the console. The field is modelArn on VectorSearchBedrockRerankingModelConfiguration. Amazon Rerank 1.0 is $1 per 1,000 queries in the calculator. Cohere Rerank 3.5 is $2 per 1,000 queries on the pricing page, with at most 100 chunks in a query. Twenty candidates stay inside one Cohere query. One hundred is the API maximum for numberOfRerankedResults.

reranked = client.retrieve(
    knowledgeBaseId=kb_id,
    retrievalQuery={"text": question},
    retrievalConfiguration={
        "vectorSearchConfiguration": {
            "numberOfResults": 20,
            "rerankingConfiguration": {
                "type": "BEDROCK_RERANKING_MODEL",
                "bedrockRerankingConfiguration": {
                    "numberOfRerankedResults": 5,
                    "modelConfiguration": {"modelArn": rerank_model_arn},
                },
            },
        }
    },
)

Generate on a smaller model, with the output cap set to the 800 tokens this example already prices. The same 10,000 calls on the calculator’s Claude Haiku 4.5 pin ($0.80 / $4.00 per million tokens) are $32 of input and $32 of output, $64 instead of $160. Keep Sonnet 5.5 for the questions that fail a review set. Nova Lite in the same calculator is $0.06 / $0.24 per million tokens, $4.32 on this token shape, and it is the wrong default for policy text until that review set passes.

boto3 bedrock-runtime Converse. model_id is the Haiku 4.5 ID from the console. The cache checkpoint sits after the retrieved context and before the question. Default TTL is 5 minutes. Set ttl to 1h only on models that document the one-hour cache. A model that only supports 5 minutes rejects the field.

runtime = boto3.client("bedrock-runtime", region_name="us-east-1")

answer = runtime.converse(
    modelId=model_id,
    system=[{"text": "Answer from the retrieved passages. Cite the passage."}],
    messages=[
        {
            "role": "user",
            "content": [
                {"text": retrieved_passages},
                {"cachePoint": {"type": "default"}},
                {"text": question},
            ],
        }
    ],
    inferenceConfig={"maxTokens": 800},
)

On 7 October 2026 the Bedrock pricing page cut Sonnet 5.5 cache-read price by 50% from the prior rate. Confirm the new dollar on the model card before you put a cache-hit percentage in a budget. Caching a context that changes every call does not hit.

One corpus, two totals

PathEmbeddingsVector storeSonnet 5.5Total
S3 Vectors$0.11$3.17$160.00$163.28
OpenSearch Classic$0.11$350.40$160.00$510.52

What broke — A console default that accepts the auto-created OpenSearch Serverless collection. Cost Explorer grouped by Bedrock feature shows on-demand inference and the reranker, and almost none of the $350.40, because those OCUs bill as OpenSearch. You see it when the Bedrock attribute total and the OpenSearch service total for the same month are different numbers. Recovery is a knowledge base on S3 Vectors, a re-ingest from the source bucket, and deletion of the Classic collection after the new index answers the same questions. Changing the sync from weekly to daily does not close the gap. Embeddings move by about two cents.

Reproduce this — Open the Bedrock Knowledge Base cost calculator. Set 5,000 documents, 16 KB average, weekly sync, 50,000 Retrieve calls, 10,000 RetrieveAndGenerate calls, Claude Sonnet 5.5, 4,000 input tokens, 800 output tokens, rerank off, Data Automation off. S3 Vectors totals $163.28 a month. OpenSearch Serverless totals $510.52. The rates are that calculator’s us-east-1 model as of 29 September 2026. The Classic 2 OCU minimum matches the OpenSearch pricing page checked 9 October 2026. The $0.24 hourly rate comes from the site OpenSearch calculator. The static pricing-page text states the 2 OCU minimum and leaves the hourly dollar in the regional price list.

Rerank on every one of the 60,000 queries adds $60 and moves the S3 total to $223.28. That is the wrong lever on a corpus whose vector store is already $3.17.

What to Do This Week

  1. In Cost Explorer, group October Bedrock spend by feature and by model. The period has to start on or after 1 September 2026.
  2. List OpenSearch Serverless collections that Bedrock created for knowledge bases. If the collection is Classic and the corpus is small, run it through the calculator before the next sync. Delete the collection only after the replacement index answers the review questions.
  3. Call Retrieve with numberOfResults of 5 and a metadata filter. Turn rerank on for the misses.
  4. Move generation to the Haiku 4.5 pin on the same token shape ($64 versus $160) and keep Sonnet 5.5 for the review-set misses.
  5. Add a budget filtered to one model. Tag the application inference profile so the attribute view and the tag view name the same application.

What This Post Doesn’t Cover

Creating the knowledge base, chunking choices, and the ingestion clicks live in the RAG setup guide. S3 Vectors product limits, including the 10,000-result query cap as of 16 June 2026, live in the S3 Vectors post. AgentCore runtime, memory, and the other eleven components live in the AgentCore pricing post.

Aurora pgvector dollars are not in the calculator. NextGen OCU behavior under a sustained query rate is not on the pricing page beyond the idle rule (scale to zero after 10 minutes). Source objects in S3 bill at S3 rates and are outside both totals above. GovCloud and the China Regions do not have the product attributes. Managed Knowledge Base agentic retrieval quality is a separate test. This post prices the call. It does not score the answers.

Frequently asked questions

When should a Knowledge Base stay on OpenSearch Serverless?
Keep OpenSearch when the query needs hybrid keyword plus vector search in one engine, or when the latency target is the one the S3 Vectors post already sends to OpenSearch or MemoryDB. Use a NextGen collection so idle compute can scale to zero after 10 minutes. Classic is the collection type with the 2 OCU minimum the calculator prices at $350.40 a month.
When should you not turn on rerank or Data Automation?
Leave rerank off when the first Retrieve pass already returns the right chunks. The calculator prices Amazon Rerank 1.0 at $1 per 1,000 queries, so 60,000 queries add $60. Leave Bedrock Data Automation off for plain text. The pricing page charges $0.010 per page for standard output when Data Automation is the Knowledge Base parser. On this corpus, weekly parsing of every document adds $12.50.
Does Cost Explorer show the whole Knowledge Base bill?
No. Product attributes cover Bedrock lines such as model, provider, inference type, and feature. OpenSearch Serverless and S3 Vectors stay on their own service lines. Compare the Bedrock attribute total with those service totals for the same month.
Is there a separate hourly fee for a Knowledge Base?
Self-managed Knowledge Bases have no separate hourly fee in the site calculator. You pay embeddings, the vector store, optional parsing and reranking, and model inference. Managed Knowledge Bases are a different price on the Bedrock pricing page: $5 per GB of raw data per month, $1 per 1,000 Retrieve calls, with the managed embedding model and managed reranker included.
What goes wrong if you group Cost Explorer before September 2026?
Product attribute data is available for time periods that start on or after 1 September 2026. An earlier start date fails with DataUnavailableException. The console control shipped on 8 October 2026.
When is a full re-ingest the right bill?
Pay it when you change the chunking strategy or the embedding model, because those changes rebuild the index. On this calculator corpus a full monthly re-embed is $0.45 of Titan tokens. Do not schedule that rebuild to save money. The OpenSearch Classic floor is $350.40 either way.
Palaniappan P
Palaniappan P

AWS Cloud Architect & AI Expert

AWS-certified cloud architect and AI expert with deep expertise in cloud migrations, cost optimization, and generative AI on AWS.

AWS ArchitectureCloud MigrationGenAI on AWSCost OptimizationDevOps

Recommended Reading

Explore All Articles »
10 min

S3 Vectors: 10,000 Results per Query (June 2026)

On June 16, 2026, S3 Vectors raised the QueryVectors limit to 10,000 results per query and cut data-processed charges up to 80% on indexes over 10M vectors. Architecture, pagination, and cost comparison vs OpenSearch and MemoryDB.