AWS Bedrock · GenAI Reference Architecture

AI Enterprise Architecture Review Board

An AI-powered platform that ingests architecture artifacts and acts like a board of specialists — Solution, Security, FinOps, Compliance, and Executive advisors — to identify risks, assess compliance, estimate cloud cost, and produce executive-ready recommendations, every answer grounded in cited evidence.

Region · ca-central-1 (Canadian data residency) Model · Claude Sonnet on Bedrock Pattern · Supervisor + 5 collaborator agents + RAG ADRs · 14 recorded decisions Infra · AWS CDK · 4 stacks · built
⎇ View on GitHub · bdot-real/arch-review ↗ 🗂️ Issues / backlog ↗
📈 ExecutiveValue, risk & cost
🏛️ ArchitectSystem design & decisions
⚙️ DeveloperComponents & contracts

What it is, in one line

Upload your architecture. Ask a question. Get a board-level review with evidence.

🤝 The business problem

Architecture review depends on scarce senior experts and tribal knowledge. Reviews are slow, inconsistent, and rarely produce an audit trail.

💡 The solution

An always-available AI review board that delivers consistent, explainable, governance-grade assessments in under a minute.

🎯 The outcome

Faster decisions, lower review cost, documented compliance evidence, and a one-page CIO summary for every engagement.

The five-specialist review board

One orchestrator coordinates five specialist perspectives — the way a real review board works.

Solution Architect

System design, patterns, scalability.

Security Architect

IAM, network, encryption, threat modeling.

FinOps

Cost estimation and optimization.

Compliance

PIPEDA, GDPR, HIPAA, EU AI Act, governance.

Executive Advisor

Business language, ROI, risk summaries.

🧭 Orchestrator

A supervisor agent that interprets the request, delegates to the right specialists in parallel, and synthesizes the final verdict — it holds no tools of its own, so it can't fabricate numbers.

What a single review delivers

Actual output from a live board run on the seeded retail-csp demo (Claude Sonnet 4.5, ca-central-1): “Can this architecture support an autonomous AI support agent handling Canadian customer PII?”

Board verdict: ⛔ DEPLOYMENT BLOCKED — overall architecture score 62/100, with 5 critical blockers to clear before processing Canadian PII. Estimated remediation: ~10 weeks to safe launch.

Architecture scorecard (from the deterministic tools)

70Security
75Reliability
29Cost efficiency
FAILCompliance

Compliance assessment

PIPEDA FAIL

No formal PII handling / data-classification standard for the AI assistant (Schedule 1 Principles 4.3 / 4.4 / 4.7).

GDPR FAIL

Same PII-governance gaps; data-residency for model inference unverified.

AI governance MISSING

No NIST AI RMF / ISO 42001 / EU AI Act classification, guardrails, or human-in-the-loop.

One-page CIO summary (generated)

📍 Current state & risks

Solid serverless foundation, but 5 critical blockers: PIPEDA gap, over-permissive IAM + no MFA, unverified data residency, no AI governance, single-region (no DR).

✅ Recommended actions

Adopt a PII standard, decompose IAM to least-privilege + enforce MFA, add Bedrock Guardrails + HITL, design multi-region DR, right-size compute.

💰 Cost & timeline

$1,016.68/mo now → $455–$673/mo optimized (34–55% savings). ~$68K one-time remediation; CIO/CISO go/no-go at the Week-10 gate.

Every figure above came from the deterministic Action Group tools + RAG retrieval with source citations (NFR-004) — not free-form model output. Full transcript: docs/demo-output.md.

Why these choices reduce business risk

🇨🇦 Data stays in Canada

Regulated data plane runs in ca-central-1, directly supporting the PIPEDA / data-residency story.

🔍 Every answer is explainable

Recommendations cite source documents with a reasoning summary and confidence score — governance-grade, not a black box.

🧾 Everything is audited

Prompt, response, sources, tools invoked, user, and timestamp are logged for every action.

💸 Pay for what you use

Serverless throughout; the demo design drops idle cost to near zero while keeping a production scale-up path.

System context & data flow

As built across four CDK stacks (Foundation · Data · Agent · Api). React SPA → API Gateway → Lambda → a Bedrock supervisor agent that delegates to five collaborator agents, each calling its own deterministic Action Groups — all on a KMS-encrypted, audited, ca-central-1 data plane.

flowchart TB U([User · CIO / Architect / Security]):::actor subgraph FE[Presentation] SPA[React SPA
Cognito SRP auth]:::fe end subgraph EDGE[Api stack] APIGW[API Gateway · v1
Cognito authorizer · throttle · validate]:::aws end subgraph APP[Api stack · Lambdas] PROJ[project / upload
retrieve]:::aws REVF[review
POST 202 + GET poll]:::aws RW[review-worker
async · 5 min]:::aws end subgraph BR[Agent stack · Bedrock] SUP[Supervisor agent
SUPERVISOR · no tools]:::agent SPEC[5 collaborator agents
SA · Sec · FinOps · Comp · Exec]:::agent FM[Claude Sonnet
cross-region inference profile]:::model end subgraph TOOLS[Lambda Action Groups · deterministic] RET[retrieveContext]:::tool REST[scoreArchitecture · estimateCost
assessCompliance · generateDiagram
generateThreatModel]:::tool end subgraph DATA[Data + Foundation stacks · ca-central-1 · KMS CMK] PT[(DynamoDB
Projects)]:::store VT[(DynamoDB
Vectors · demo)]:::store AT[(DynamoDB
Audit · append-only)]:::store S3U[(S3 uploads)]:::store S3R[(S3 reports + artifacts)]:::store OSS[(OpenSearch Serverless
Bedrock KB · prod mode)]:::store end ING[ingest Lambda
extract · chunk · Titan embed]:::aws EB{{EventBridge
Object Created}}:::aws OBS[CloudWatch
dashboard · alarms · X-Ray]:::sec BUD[Budgets → SNS]:::sec U --> SPA --> APIGW APIGW --> PROJ & REVF SPA -. pre-signed PUT .-> S3U REVF -. InvocationType Event .-> RW --> SUP SUP --> SPEC --> FM SPEC --> RET & REST RET --> VT RET -. prod .-> OSS S3U --> EB --> ING --> VT S3U -. prod ingest .-> OSS RW --> S3R PROJ --> PT PROJ & REVF & RW --> AT APIGW --> OBS BUD -.cost guardrail.-> DATA classDef actor fill:#1b2440,stroke:#3a4d77,color:#e8edf7; classDef fe fill:#13233f,stroke:#4f8cff,color:#dbe7ff; classDef aws fill:#241a0a,stroke:#ff9900,color:#ffe9c7; classDef agent fill:#1d1530,stroke:#b18cff,color:#eaddff; classDef model fill:#102a22,stroke:#37d399,color:#c9ffe9; classDef tool fill:#10203a,stroke:#4f8cff,color:#d6e6ff; classDef store fill:#0f1a30,stroke:#33476f,color:#cdd9f0; classDef sec fill:#2a1320,stroke:#ff6b6b,color:#ffd5d9;
AWS managed service Bedrock agent Foundation model Action Group / tool Security & governance

Multi-agent collaboration topology

ADR-0013, as built: a SUPERVISOR orchestrator that owns no tools — it plans, delegates, and synthesises — with five separate collaborator agents, each granted only the Action Groups it needs (least privilege per specialist). All specialists share retrieveContext for grounding.

flowchart LR ORC[Supervisor / Orchestrator
plans · delegates · synthesises]:::sup SA[Solution Architect]:::ag SEC[Security Architect]:::ag FIN[FinOps]:::ag CMP[Compliance]:::ag EXE[Executive Advisor]:::ag ORC --> SA & SEC & FIN & CMP & EXE RC([retrieveContext]):::t SC([scoreArchitecture]):::t DG([generateDiagram]):::t TM([generateThreatModel]):::t EC([estimateCost]):::t AC([assessCompliance]):::t SA --> RC & SC & DG SEC --> RC & SC & TM FIN --> RC & EC CMP --> RC & AC EXE --> RC & SC & EC classDef sup fill:#241a0a,stroke:#ff9900,color:#ffe9c7; classDef ag fill:#1d1530,stroke:#b18cff,color:#eaddff; classDef t fill:#10203a,stroke:#4f8cff,color:#d6e6ff;
Why it matters: a specialist can invoke only its own tools (e.g. FinOps cannot run the threat model), so least-privilege is enforced at the agent boundary — and the supervisor can never fabricate a number because it holds no tools at all.

Review request lifecycle — async job + poll

A full board run takes ~30–60s, past API Gateway's hard 29s timeout — so POST /review returns 202 { reviewId } immediately and a worker Lambda runs the agent off the request (ADR-0012).

sequenceDiagram autonumber participant U as SPA (api.review) participant G as API Gateway participant V as review Lambda participant J as Projects table (job) participant W as review-worker (async) participant S as Supervisor agent participant C as Collaborator agents participant T as Action Groups participant R as Reports S3 U->>G: POST /review (Cognito auth) G->>V: invoke (validated body) V->>J: create job REVIEW#id · status running V-->>W: invoke async (InvocationType Event) V-->>U: 202 { reviewId } Note over S,C: Supervisor plans, delegates in parallel W->>S: InvokeAgent(question, projectId) S->>C: delegate to needed specialists C->>T: retrieveContext + scoring/cost/compliance/threat T-->>C: deterministic JSON + citations C-->>S: evidence-backed findings S-->>W: synthesised recommendation W->>R: write markdown report artifact W->>J: status done + audit record loop poll until done/failed U->>G: GET /review?reviewId G->>V: read job V-->>U: status (+ result when done) end
Key principle: the LLM orchestrates and explains; it never invents numbers. Scores, costs, and compliance verdicts come from deterministic, unit-tested Lambda Action Groups — satisfying explainability (NFR-004). The worker writes both the report artifact and the append-only review.invoke audit record (NFR-003).

Architecture Decision Records

Every significant decision is captured as an immutable, Nygard-style ADR in docs/adr/.

ADRDecisionRationale (short)Status
0001Record architecture decisionsModel the governance discipline the product enforces.Accepted
0002Bedrock Agents for orchestrationManaged, AWS-native multi-agent coordination.Accepted
0003Bedrock KB + OpenSearch Serverless (RAG)Managed retrieval with citations — production target.Accepted*
0004Claude Sonnet as foundation modelStrong long-doc reasoning + reliable tool calling.Accepted
0005Lambda Action Groups for toolsDeterministic, testable, auditable capabilities.Accepted
0006Cognito + RBAC + KMSManaged identity, persona RBAC, CMK encryption.Accepted
0007S3 for documents & artifactsDurable object store + event-driven ingestion.Accepted
0008React SPA behind API GatewayClean FE/BE split; central auth & throttling.Accepted
0009IaC tooling — AWS CDK (TypeScript)One language across FE + infra; built as 4 stacks.Accepted
0010ca-central-1 regionCanadian data residency for PIPEDA.Accepted
0011Lightweight vector store (demo)Near-zero idle cost; RAG as an Action Group.Accepted
0012Async review via job + pollSurvives ~30–60s board runs past API Gateway's 29s limit.Accepted
0013Supervisor + collaborator topologySeparate agents, least-privilege tools per specialist.Accepted†
0014EventBridge ingestion triggerAuto-ingest on upload with no cross-stack cycle.Accepted

* ADR-0003 remains the production target; ADR-0011 amends it to scope OpenSearch Serverless to production and introduce the demo retrieval path.  † ADR-0013 refines ADR-0002 with the as-built supervisor/collaborator wiring (requires aws-cdk-lib ≥ 2.260).

Two retrieval architectures, one contract

Demo and production differ only behind the retrieveContext interface — the agent and UI are retrieval-backend agnostic.

Demo path — lightweight (built)

Upload → S3 → EventBridge → ingest extracts/chunks/embeds (Titan) → vectors + citations in a DynamoDB table → retrieveContext Lambda runs similarity search and returns top-k chunks + citations.

  • Near-zero idle cost (no OpenSearch OCUs)
  • RAG is an explicit, controllable Action Group
  • Scale/concurrency limited — demo corpus only

Production path — managed

Bedrock Knowledge Base backed by OpenSearch Serverless (built as OpenSearchKnowledgeBase, switched on by retrievalMode: 'prod'): managed chunking/embedding/sync and native Agent↔KB integration with horizontal scale.

  • Always-on OCU baseline cost
  • Managed sync & scale
  • Same downstream contract: top-k chunks + citations

Cross-cutting concerns

🔐 Security

Cognito SRP authN, RBAC by persona (enforced, #45), KMS CMKs at rest, TLS in transit, least-privilege IAM.

🧾 Auditability

Append-only DynamoDB audit table — writers hold PutItem only; prompt, response, sources, actions, user, timestamp.

🔎 Explainability

Evidence, source docs, reasoning summary, confidence on every recommendation.

📊 Observability

CloudWatch dashboard + alarms (API 5xx, worker errors), X-Ray tracing, Budgets→SNS cost alerts (#11/#16/#19).

Component map — by CDK stack

As built in infra/: four stacks — Foundation (KMS, S3, Cognito, Budgets), Data (DynamoDB, ingest, EventBridge), Agent (supervisor + 5 collaborators + Action Groups), Api (API Gateway + handlers + observability).

flowchart LR subgraph FE[Frontend] R[React SPA
Tailwind · mock mode]:::fe end subgraph APIS[Api stack] GW[API Gateway v1
Cognito authz · throttle]:::aws PRJ[project · upload
retrieve]:::aws REV[review
POST/GET]:::aws RW[review-worker
async 5 min]:::aws DASH[CloudWatch
dashboard + alarms]:::sec end subgraph FND[Foundation stack] CG[Cognito + RBAC groups]:::sec KMS[KMS CMK]:::sec B1[(S3 uploads)]:::store B2[(S3 reports)]:::store B3[(S3 artifacts)]:::store BUD[Budgets → SNS]:::sec end subgraph DAT[Data stack] PT[(DynamoDB Projects)]:::store VT[(DynamoDB Vectors)]:::store AT[(DynamoDB Audit)]:::store ING[ingest Lambda]:::aws EB{{EventBridge rule}}:::aws end subgraph AGT[Agent stack] ORC[Supervisor agent]:::agent SP[5 collaborator agents]:::agent CS[Claude Sonnet]:::model A1[retrieveContext]:::tool AX[score · cost · compliance
diagram · threatModel]:::tool end R-->GW CG-.authz.->GW GW-->PRJ & REV R-.presign PUT.->B1 REV-.async.->RW-->ORC-->SP-->CS SP-->A1-->VT SP-->AX-->PT B1-->EB-->ING-->VT RW-->B2 AX-->B3 PRJ-->PT PRJ & REV & RW-->AT GW-->DASH classDef fe fill:#13233f,stroke:#4f8cff,color:#dbe7ff; classDef aws fill:#241a0a,stroke:#ff9900,color:#ffe9c7; classDef agent fill:#1d1530,stroke:#b18cff,color:#eaddff; classDef model fill:#102a22,stroke:#37d399,color:#c9ffe9; classDef tool fill:#10203a,stroke:#4f8cff,color:#d6e6ff; classDef store fill:#0f1a30,stroke:#33476f,color:#cdd9f0; classDef sec fill:#2a1320,stroke:#ff6b6b,color:#ffd5d9;

Action Group contracts

Typed JSON request/response per tool. Business logic lives in Lambda (unit-testable), not in prompts.

Action GroupFunctionOutput shape
AG-001 ScoringscoreArchitecture(){ security, reliability, cost, compliance } (0–100)
AG-002 CostestimateCost()Monthly estimate across EC2/ECS/Lambda/OpenSearch/Bedrock/S3/RDS
AG-003 ComplianceassessCompliance()pass / warning / fail per PIPEDA · GDPR · HIPAA · SOC2 · ISO27001
AG-004 DiagramgenerateDiagram()Mermaid or PlantUML source → S3 artifacts bucket
AG-005 Threat modelgenerateThreatModel()STRIDE analysis · risk matrix · mitigations → S3 artifacts
RAG RetrievalretrieveContext()top-k chunks + source citations (query, projectId, topK)
Sample — AG-001: request {"architectureId":"123"} → response {"security":87,"reliability":91,"cost":73,"compliance":80}

Which specialist gets which tools (least privilege, ADR-0013)

Collaborator agentGranted Action Groups
Solution ArchitectretrieveContext · scoreArchitecture · generateDiagram
Security ArchitectretrieveContext · scoreArchitecture · generateThreatModel
FinOpsretrieveContext · estimateCost
ComplianceretrieveContext · assessCompliance
Executive AdvisorretrieveContext · scoreArchitecture · estimateCost
Supervisor / Orchestratornone — plans, delegates, synthesises only

Ingestion pipeline

Event-driven: an S3 upload triggers chunking, embedding, and indexing.

1

Upload — SPA gets a pre-signed URL and PUTs the file (PDF/DOCX/TXT/MD/PNG/JPG/SVG/PPTX) directly to the KMS-encrypted uploads bucket.

2

Trigger — the bucket emits an Object Created event to EventBridge; a rule in the Data stack (matched by bucket name, so no cross-stack cycle — ADR-0014) targets the ingest Lambda.

3

Extract — txt/md read directly; PDF/DOCX/PPTX go through the extractor (#22).

4

Chunk + embed — text is chunked and embedded with amazon.titan-embed-text.

5

Index — demo: write vectors + source citations to the DynamoDB Vectors table. Prod (retrievalMode: 'prod'): sync into OpenSearch Serverless via the Bedrock KB.

6

Retrieve — at query time retrieveContext embeds the query, runs similarity search, returns top-k chunks + citations.

Tech stack & build notes

Frontend

React SPA (Tailwind), Cognito SRP auth, REST via API Gateway, direct-to-S3 pre-signed uploads. Local mock mode + Playwright e2e (#71/#73).

Backend

NodejsFunction Lambdas: project, upload, retrieve, review + review-worker (async). Job + poll for the <60s budget (ADR-0012).

AI

Bedrock supervisor + 5 collaborator agents, Claude Sonnet via cross-region inference profile; Titan embeddings for RAG.

Storage

3 S3 buckets (uploads / reports / artifacts) + 3 DynamoDB tables (Projects / Vectors / Audit), all KMS-CMK encrypted, PITR on.

Security

Cognito + persona RBAC (enforced), KMS CMKs, append-only audit (PutItem-only), least-privilege IAM per Lambda & per agent.

IaC & ops

AWS CDK (TypeScript), 4 stacks, Jest construct tests; CloudWatch dashboard + alarms, X-Ray, Budgets→SNS; pinned to ca-central-1.

Things to watch (known risks)

  • Shared model RPM under fan-out — parallel collaborators share the bound model's RPM; a low cross-region quota causes 429s during a board run (ADR-0013 / NFR-001). Quota increase + retries.
  • Pricing drift — estimateCost needs a maintained AWS pricing source (AG-002).
  • Two retrieval code paths — demo DynamoDB store vs prod OpenSearch KB; mitigated by the shared retrieveContext contract.
  • Open CORS — uploads bucket and API currently allow all origins; must be restricted to the SPA origin before prod (TODOs in foundation-stack.ts / api-stack.ts, ADR-0008).
  • Model availability in ca-central-1 — Claude Sonnet is reached via a cross-region inference profile; the residency exception needs compliance sign-off (ADR-0010).
  • L1 Bedrock constructs — verbose CfnAgent wiring (needs aws-cdk-lib ≥ 2.260); L2 aws-bedrock-alpha is the eventual target (ADR-0009/0013).

Explore the repository

Jump straight to the code on GitHub — bdot-real/arch-review.

Infrastructure (AWS CDK)

Frontend, docs & ops