Interactive architecture and concept diagrams spanning the Score Anything Loyalty platform I build at 3 Halves Labs, my independent, openly-built projects (Sepulchre, OSSS and more), and reference studies. Many have an Executive / Developer toggle, and there's a deliberate focus on AI architectures that make different methods and decisions — hosted multi-agent systems, on-box constrained-decoding models, and computer-vision and physics-based estimation — pulled together in the AI Architecture Decisions capstone.
The pieces below live in different tracks, but they answer one question in different ways: what kind of AI does the job, given what matters most? This capstone puts them side by side.
A production retrieval-augmented-generation pipeline drawn as a schematic node-graph: the offline indexing chain (sources → chunk → embed → governed vector index) and the online query chain (query → guardrail → embed → retrieve → rerank → assemble → generate), meeting at retrieval where access controls filter what can be pulled, with an evaluation loop closing back on the index. Filter by path; tap any node.
One governed agent request as a sequence diagram across six lifelines (user, gateway, agent, retrieval, tools, LLM): identity in at the top, a reason–act loop in the middle where the model calls tools and reads observations, a guarded and cited answer out the bottom, and every message on the immutable audit log. Play it, step it, or tap any message.
The high-detail path from raw sources to training-ready datasets: source & collect, ingest & parse, deduplicate (exact + fuzzy MinHash/LSH), quality filter, safety & PII (hard gate), decontaminate against benchmarks (hard gate), mix & weight, tokenize & pack, and version & register — plus the curation deep-dive, data mixing and tokenization, and post-training data (SFT, RLHF/preference, eval/red-team, synthetic with a model-collapse caution).
The accountability layer over the pipeline: provenance and a licence taxonomy (permissive / attribution / restricted / prohibited), privacy and individual rights (de-identification, consent, opt-out and an honest take on right-to-be-forgotten limits), fairness and integrity (representativeness, documented safety, benchmark decontamination, synthetic-data governance), and lineage — a data bill of materials linking every model checkpoint back through its dataset version to the original sources.
How a trained model becomes a fast, reliable API: the serving path (gateway → router → scheduler → continuous batching → KV cache → GPU workers → decoding → streaming → observability), the throughput and latency levers (continuous batching, paged attention, prefix caching, speculative decoding, quantization, parallelism) with the metrics that matter (TTFT, TPOT, throughput, concurrency), and scaling and reliability (autoscaling, fallback, canary, cost controls).
The safety layer around an LLM app: input guards before the model (auth/rate-limit, prompt-injection and jailbreak detection, PII redaction, topic/policy filter, input validation), output guards after it (safety/toxicity, PII/secret-leak scan, groundedness, citation enforcement, schema validation, policy compliance), and a guarded-request walk showing least-privilege tools and human-in-the-loop around the loop — six of eight steps are enforcement points.
The model lifecycle drawn as a loop rather than a line: data & features → train → evaluate (gate) → register → deploy (gate) → serve → monitor → retrain trigger, arranged around a ring with the two hard gates marked, a walkable rotation, and per-stage detail (inputs, outputs, and what stops the cycle). New diagram type — a radial cycle.
A supervisor pattern as a hierarchy: one orchestrator plans and routes sub-tasks to four specialist agents (researcher, analyst, coder, writer/critic), who all draw on the same governed capability bus — tools, retrieval, memory and models — with a guardrail frame around every model and tool call and a human-in-the-loop approval gate. The peer/handoff and reflection patterns sit alongside as context. New diagram type — a hierarchy tree.
The sports-loyalty platform end to end: how a signal becomes points, the data and AI that sit behind the loyalty markets, and the governance and compliance around them.
The full journey of a browser signal becoming loyalty points: capture → consent-gated enqueue → batching → the wire envelope → middleware queue → ERP rules → points. Includes the lifecycle, a system map, a subsystem catalogue and the contract & failure-mode reference.
The reference architecture between the edge and the models: a durable streaming bus, stream and batch processing, a bronze/silver/gold lakehouse, a feature store and activation — with cross-cutting governance, lineage, quality and SLOs. Every block names the industry archetype it maps to.
The map of the AI estate: where models and agents live, what data each touches, and the governance that binds them. The consumer end of the platform, sitting on the data spine and inside the guardrails.
The governance layer that spans the whole platform: consent capture and enforcement, data residency, lineage, and the access policy that decides who and what can read each class of data.
The federation topology seen through analytical lenses — risk heatmap, blast-radius cascade, data classification, sovereignty and latency — plus the brokered OIDC flow and the embedded third-party SDK. Click any node or edge to open its risk register.
A user is up to four records across four systems, linked by shared ids. Four entry paths — only two create a user — with a step-through of the primary lazy, JWT-triggered Auth0 → Mongo → RabbitMQ → Odoo flow, the cross-system identity trail, and the caveats the trace surfaced.
Fire an event and watch the engine match it against rules (type, url, selector, threshold, payload), check limits (once / cooldown / per-match), award karma, advance challenges and goals, and move the player up a tier. The business core, picking up exactly where the Tracking SDK hands off.
The economics beneath the loyalty loop: set a season budget ceiling, tune base earn rates and per-action levers (fixed and spend-based), and watch the earn schedule that results. Then trace where the points land as liability, float, redemption and breakage, so issuance stays inside the budget. Exports to CSV.
The integrity layer beneath the engine: an append-only, double-entry, idempotent and reversible ledger that finance can reconcile and audit. Post entries, watch the books stay balanced, try a duplicate and see exactly-once reject it, then break and rebuild the reconciliation.
Defense-in-depth for Score AI as an MCP server, seen through three lenses (Security, Privacy, Executive). A request-trace simulator walks a request through all eight guardrail layers, an access evaluator and an egress simulator block over-reach and leakage, plus a threat model (OWASP LLM / MCP / ATLAS) and a conformance matrix to ISO 42001 / 27001 / 27701, SOC 2, GDPR and the EU AI Act.
The prediction service behind the loyalty markets, seen as executive or engineer. A supervisor orchestrates specialist predictor agents; an ensemble/critic/risk chain blends, calibrates and guards them; a market-maker prices the result; and a closed learning loop grades every outcome. Sits on the data platform, governed by the AI guardrails, exposed via the Score AI MCP.
What Score has to put in place for a SOC 2, mapped to its actual stack. The five Trust Services Criteria with a scope selector, the nine Common Criteria (CC1–CC9) and the controls that satisfy them, the practical control domains, the road from gap assessment to a Type II report, and a 36-item readiness checklist.
The markerless pipeline in nine clickable stages: MoveNet pose estimation, camera-view classification and normalization, Newton-Euler inverse dynamics for forces and joint torques, bounds validation, and elite comparison — plus the physics table, the running metrics and injury-risk indicators, and the lab/simulation/reliability validation.
A different class of AI architecture: probabilistic state estimation, not language models. The complementary pose/IMU sources, a cyclable Extended Kalman Filter predict/update loop, and the synchronization (cross-correlation + confidence weighting) and graceful degradation that keep it robust when a sensor drops out.
Projects I build and govern in the open, outside 3HL: a zero-knowledge credential broker, an open scheduling standard, an ISO 20022 MCP server, and an AI architecture-review tool.
Built on Vault and vendor-neutral by design, where the operator's inability to read tenant secrets is enforced by policy and provable from the audit log — with an on-box AI substrate so its AI features never ship data to a third party.
Start here for the whole picture: the standalone Sepulchre architecture site brings the system design, the zero-knowledge guarantee and the on-box AI substrate together in one place, with the reasoning behind each. The individual diagrams below drill into the pieces.
The v2 architecture: a stateless broker over a Vault cryptographic core, Postgres for metadata only, a signed audit pipeline, and pluggable roots of trust. Five actors, pooled and siloed tenancy from one binary, the two canonical data flows, and the deployment topologies with a bounded-blast-radius failure model.
How the operator's inability to read tenant secrets is enforced and proven. Walk an attempted operator read through the policy, code, CI and operational layers that each stop it independently, then the per-tenant key hierarchy, the signed hash-chained audit event and the continuous attestation that makes the empty human-read list provable. BYOK for the strongest form.
Audit Q&A and compliance narratives without a hosted API: an in-process small language model (Qwen2.5-1.5B via node-llama-cpp) that never sees plaintext, runs in a credential-less worker, and is grammar-constrained (GBNF) to emit only a whitelisted audit-query object — never free-form SQL. The architecture, an interactive constrained-decoding walkthrough, the decision record and the benchmark plan.
The Open Sports Scheduling Standard in the open: the spec modules and JSON schemas, the constraint and objective registries, the conformance suite and profiles. Read the spec, clone it, or open an issue — this is where the standard is governed.
A machine-readable language for sports scheduling that standardizes the problem, not the solver: the interop model and spec modules, the core data model (teams, venues physical and virtual, fixtures, officials, resources), a 51-constraint registry across 11 categories with six penalty families, and explainable, auditable results with Pareto alternatives, profiles and a conformance suite.
The source behind the AI Enterprise Architecture Review Board: the Bedrock multi-agent orchestrator and specialist reviewers, the CDK stacks and the React front end. Clone it, read the code, or run it yourself.
Upload an architecture, ask a question, get a board-level review with evidence. One orchestrator coordinates five specialist perspectives the way a real review board works, built across four CDK stacks (Foundation · Data · Agent · Api): a React SPA through API Gateway and Lambda to Amazon Bedrock.
Reference and solution architectures, plus interactive tools — some designed from scratch for a platform or a role, some built to study a published design.
A reference architecture for conversational and autonomous-agent AI on the Azure Databricks Lakehouse: ingestion → medallion → Unity Catalog governance → serving, the chat/voice/agent layer with governed retrieval and tools (plus Vertex AI / Gemini as a multi-cloud model option), six security pillars (Entra ID, Unity Catalog row/column security, Private Link, customer-managed keys, Mosaic AI Gateway guardrails, audit), a quality layer (data-quality tests, agent evaluation, a CI gate, and Lakehouse Monitoring for data and model drift), and a governed request walk. The through-line: the AI inherits the platform's access controls, so it can never surface data the user isn't entitled to.
The deployment and scaling companion to the secure-AI architecture: the Azure landing zone (control-plane vs data-plane split, injected VNet with host/container subnets, front-end and back-end Private Link, ADLS and Key Vault behind private endpoints, controlled egress), compute scaling (serverless vs classic per workload, model-serving concurrency, Vector Search sizing), and a cost model showing what each component's bill is proportional to and which lever bends it — including the monitoring and evaluation costs that scale with traffic.
A clean, interactive rebuild of AWS's "Open Banking on AWS" reference architecture: nine zones from consumer through edge, API Gateway (mTLS + OAuth 2.0 + TSP), ECS/Fargate microservices and the hybrid path to the bank core, all 18 numbered components clickable, plus Account-Information and Payment request-flow walkthroughs. Architecture and service names © Amazon Web Services.
Pick a caching strategy, fire a read, and watch data move between the component, the client cache and the origin — with freshness, latency and offline outcomes. Includes a comparison matrix and a mapping to Cache-Control, TanStack Query, SWR, Apollo and Workbox.