Illustrative decision records

Same system. Same day. Different authority.

Decision record / A

executed
Event
storage node unresponsive
Diagnosis
logs reviewed, memory exhaustion, no data at risk
Action
failover
Required 0.75 Scored 0.94

cleared the bar, no human involved

Decision record / B

escalated
Event
anomalous ingress traffic
Diagnosis
source range flagged, blast radius includes production
Action
firewall rule change
Required 0.97 Scored 0.88

diagnosis done, a person decides

Production AI / security governance

AI systems for companies that can't afford to get it wrong.

Accrava builds production AI, and the governance that lets it pass a security review. Most teams have one of those. You need both.

Start a conversation

Engagements / 02

Engagements

Operations / after launch

Built for the operating model you need.

Runtime
Deploy into your environment or Accrava's private cloud.
Ownership
Hand off with documentation and runbooks, or keep Accrava operating and maintaining the agents.
Models
Support major LLM providers with routing and controls that do not depend on a single model.

Building the system and proving its controls work are usually treated as separate jobs. That separation is where AI projects stall. Accrava brings production implementation, risk-based controls, and evidence that holds up under review into the same engagement.

Fit check / before engagement

Who this work is for

A good fit if

  • You have AI in production, or you are about to, and the risk is real rather than theoretical
  • You operate under regulatory, contractual, or customer-imposed security requirements
  • Your security team is small and covering more than it should
  • You want someone who will write the code and also sign the assessment

Not a good fit if

  • You want a strategy deck and a roadmap with no implementation
  • You want the cheapest available builder
  • You want someone to certify a system they had no hand in reviewing properly

Selected work / inspect the record

The work

01

Autonomous security and infrastructure operations

One person covering an enterprise-scale platform, with agents handling the work that did not need a human.

#agents

Problem

The security function was one person, covering a Kubernetes platform serving hundreds of millions of requests a day. Every alert required someone to pull the relevant logs, decide whether it was real, and act. That work competed directly with everything else on one person's plate, and response time on genuinely actionable events depended on somebody being available to look.

Built

  • Agents that ingested security and infrastructure telemetry, correlated it against the relevant logs, assigned severity, and routed
  • Low-risk remediation such as service restarts and failovers executed autonomously once the agent had reviewed the logs
  • Higher-risk changes such as firewall rule modifications were diagnosed by the agent, which produced an impact and risk assessment and presented it to the team as a single approve or decline decision, then carried out the change on approval

Result

Nobody read logs to determine whether an alert was real. Routine remediation happened without a person. Human judgment stayed on the changes that warranted it, but the diagnostic work in front of that decision was already done and documented. One person covered an environment that would normally require a team.

02

RAG platform for support and engineering

Documentation that answered its own questions, for customers and engineers alike.

#rag

Problem

Documentation drifted out of date, and the same questions arrived repeatedly from both customers and internal engineers. Answering them consumed engineering time that should have gone to shipping.

Built

  • A RAG platform over product documentation and internal knowledge, serving customers and engineers from the same corpus
  • Continuous daily reingestion so answers reflected the current state of the system rather than a snapshot taken at launch
  • Sentiment analysis on incoming requests, and context-aware code assistance for the engineering side

Result

80% of tickets were resolved end to end with no human in the loop. Fully handled and closed, not deflected to a help article. Engineers got that time back.

03

Multi-model governance and confidence gating

The mechanism that decided what an agent was allowed to do on its own.

#governance

Problem

Running AI against live infrastructure raises a question that had no standard answer in 2022. How do you know the output is reliable enough to act on? A single model is a single point of failure twice over. It goes down, and it gets confidently wrong. Neither is acceptable when the output changes a production system.

Built

  • Routing across Anthropic, OpenAI, and Google models based on which was strongest for the task, which had the context window the job needed, and which was actually available, with automatic failover during provider outages
  • Independent cross-model review, where a different model evaluated the output rather than trusting a model's assessment of its own work
  • A confidence score combining the model's own assessment with that independent review, gated against thresholds that scaled with risk. A firewall rule change, or anything that would normally require admin privileges, had to clear a far higher bar than classifying the sentiment of a support ticket
  • A local persistent state store so agents could resume prior work across sessions, built before any provider offered the equivalent

Result

The autonomy tiers in the operations work had a real gate behind them instead of a policy on paper. Provider outages stopped being an operational event. Every autonomous action carried a score and a reviewing model's assessment, which meant the whole system was auditable after the fact.

Code / available for inspection Open-source work on GitHub
Chris Garcia, founder of Accrava

Chris Garcia / Founder

Founder / the experience behind Accrava

About Accrava

Focus
Production AI and AI governance for regulated and security-sensitive operations
Credentials
CISSP / CISM / MS, Cybersecurity
AI systems
Deploying production AI agents since late 2022
Enterprise leadership
Nearly 14 years at a Fortune 1000 manufacturer, including as Director of Network and Security. Led global teams responsible for infrastructure across 130+ sites, a budget over $10M, and 25+ M&A technology integrations.
Production infrastructure
Built the infrastructure and security foundation for a novel storage platform as employee number one and founding security executive. Designed the Kubernetes, Ansible, and Docker stack behind hundreds of millions of daily requests.

I have spent the last few years researching and building production AI with governance and security designed in from the start. That work included model routing, independent review, persistent memory, and risk-based approval gates before those patterns were available as products. Accrava brings that experience to systems that need to work in production and hold up under review.

Chris Garcia / Founder

Contact / project inquiry

Start a conversation