Blog
ai-agentcoding-agentair-gappedopen-sourceplatform-engineeringai-governance

Inside IBM Bob (Part 2) — Designing an Open-Source Coding Agent Platform for Air-Gapped Environments

We mapped IBM Bob's components to open-source equivalents one by one to design a development agent platform that can run in an air-gapped environment with no external connectivity. It covers layer-by-layer design — harness, sandbox, model gateway, policy, modernization pipeline, observability, and supply chain — along with configuration examples and a phased roadmap.

Data DynamicsOctober 4, 202622 min read

In Part 1, we analyzed IBM Bob. In short, Bob's value lies less in the model itself than in its common harness, role-based modes and approval policies, auditing, modernization workflows, and a deployment model that brings all of this inside the customer's air-gapped network.

This post covers how we would design the same platform if we built it ourselves with open source. The goal is not to clone Bob as-is, but to use the structure Bob demonstrates as a baseline while addressing the weaknesses identified in Part 1 (no system-level sandbox, limited model choice in air-gapped environments).

This post is a design document. The configuration examples are meant to illustrate the structure and should be validated against each project's documentation before being used in production. Licenses and project status for the open-source projects are as of October 4, 2026, and may change.

1. Requirements

We start by laying out the requirements an air-gapped development agent platform must meet.

AreaRequirement
DataSource code, prompts, and logs never leave the organization's boundary
IsolationCommands the agent runs execute only inside an isolated environment and cannot communicate with anything other than approved internal services
PermissionsThe tools available and the files that can be modified differ by role (ask, plan, implement, review)
ApprovalState-changing actions are judged by policy as auto-allowed, approval-required, or blocked
AuditingIt is possible to reconstruct who (which agent) did what, and on what basis
CostUsage can be measured per team and project, with caps applied
ModelsNot tied to any specific model; models can be switched by task type
SurfacesUsable both in developers' editors and in CI pipelines
ModernizationLarge-scale changes such as framework and JDK upgrades are carried out as repeatable procedures
OperationsModels, tools, and vulnerability databases can be updated without external connectivity

2. Design principles

We condensed the requirements into five principles. Every layer design that follows adheres to them.

  1. Separate layers with standard protocols. ACP between editors and agents, MCP between agents and tools, an OpenAI-compatible API between agents and models, and AGENTS.md for rules. When the boundaries between layers are standard, if an open-source project in one layer stalls, you only need to replace that layer.
  2. Isolate outside the agent. An agent's internal ignore lists and approval settings are conveniences, not security boundaries. The security boundary is the responsibility of the container runtime and network policies.
  3. Hide models behind a gateway. The harness knows only role aliases such as agent-plan and agent-code, not actual model names. Model swaps, routing, budgets, and logging are handled at the gateway.
  4. Deterministic tools first, models for the rest. Whatever can be done with compilers, tests, static analysis, and code transformation tools should be done with them; models are used to interpret those results and fix the remaining problems.
  5. Manage policies and configuration in Git. Mode definitions, approval policies, rules, and gateway configuration all live in a repository and are deployed after review. Policy changes are subject to auditing too.

3. Overall architecture

Loading diagram…

Compared with Bob's structure from Part 1, the biggest difference is that Bob's "common harness" is split into ② and ③, separating agent execution from the command execution environment.

4. Mapping Bob's components to open source

Bob componentOpen-source candidatesLicenseWhat to check when choosing
Harness + BobShellOpenAI Codex CLI, OpenCode, GooseApache-2.0, MIT, Apache-2.0Non-interactive execution, custom model endpoints, MCP and ACP support, approval hooks
Bob IDEcode-server + Cline or ContinueMIT, Apache-2.0The extension's mode and approval features, air-gapped installation (bringing in extension files)
Custom mode formatRoo Code–family extensionsApache-2.0Roo Code was archived in May 2026. Check the maintenance status of community forks
Editor integrationAgent Client Protocol (ACP)Apache-2.0Confirm ACP support across all harness candidates
Tool integrationModel Context Protocol (MCP) serversVaries by serverAuthentication and permission scope, audit logs
Model routingLiteLLMMIT (some enterprise features separate)Whether the features you need (virtual keys, budgets, SSO) fall within the open-source scope
Model servingvLLM, SGLangApache-2.0Tool-call parser support, long-context performance
ModelsOpen-weight coding models (Qwen coder family, Mistral Devstral family, OpenAI gpt-oss, NVIDIA Nemotron, etc.)Varies by modelTerms for commercial use, tool-calling quality, results on your internal evaluation set
Code understandingtree-sitter, language servers (LSP), pgvector · QdrantMIT, varies by server, PostgreSQL · Apache-2.0Indexing time for large repositories, respecting access permissions
Java modernization packageOpenRewrite, Konveyor (including Kai)Core Apache-2.0 (some recipe modules under the Moderne Source Available License), Apache-2.0Per-module recipe licenses — some modules can be applied internally but cannot be resold in commercial products
Inline security checksSemgrep CE, gitleaks, TrivyLGPL-2.1, MIT, Apache-2.0Offline updates for rules and vulnerability databases
SandboxKubernetes/OpenShift + gVisor or Kata ContainersApache-2.0Runtime compatibility, build tool performance
Approval policyOpen Policy Agent (OPA)Apache-2.0Integration point with the harness
BobalyticsOpenTelemetry, Langfuse, GrafanaApache-2.0, MIT (core), AGPL-3.0Grafana is AGPL — internal use is generally fine, but check your policy
Air-gapped deploymentHarbor (images), package mirrors (Nexus Repository, etc.)Apache-2.0, varies by productUsage terms for the mirror product's free edition

There is effectively no mature open-source equivalent to the IBM i (RPG) and mainframe (Z) packages. We explicitly leave this area as a limitation of the open-source approach.

5. Layer-by-layer design

5.1 User surfaces: ACP for editors, non-interactive for CI

There are three user surfaces.

  • Editors: Editors that support ACP (Zed, JetBrains, Neovim, etc.) connect by launching the harness CLI as an ACP server. The editor renders only the UI and approval dialogs, while the harness handles execution. This is the same structure Bob uses with bob acp.
  • Browser IDE: In air-gapped environments where installing tools on developer PCs is difficult, run code-server internally and distribute an image with the agent extension preinstalled. Extension files (.vsix) are brought in through the supply chain procedure (5.9).
  • CI: Repetitive tasks (dependency upgrades, documentation updates, draft reviews) run in pipelines using the harness's non-interactive mode, and results are always left as PRs.

Standardize on one harness, but keep it replaceable. The comparison criteria are as follows.

CriterionHow to check
Custom model endpointsCan you configure the OpenAI-compatible gateway address and role aliases?
Approval hooksCan it call an external policy engine before executing a tool? (If not, a wrapper is needed)
Rules fileDoes it read AGENTS.md?
Non-interactive executionDoes it reliably produce exit codes and artifacts in CI?
ObservabilityDoes it export trace data via OpenTelemetry?
Project healthRelease cadence, number of contributors, governance (foundation membership)

The last criterion has a real-world precedent. Roo Code, which used a format nearly identical to Bob's mode format, had its repository archived by the original development team in May 2026. This is why principle 1 — separating layers with standards — matters.

5.2 Harness: modes, rules, and approvals as code

Modes. Since every harness has a different mode configuration format, the platform keeps its source definitions in a neutral format in Git and converts them into each harness's format for deployment. Here is an example source definition modeled on Bob's format.

# platform-policy/modes.yaml — platform standard modes (source definition)
modes:
  - slug: ask
    description: Code analysis and Q&A. Changes nothing.
    tools: [read, code_search]
  - slug: plan
    description: Change planning. Writes plan documents only.
    tools: [read, code_search, edit]
    edit_paths: ["docs/plans/**"]
  - slug: code
    description: Implementation and testing.
    tools: [read, code_search, edit, execute, mcp]
    edit_paths: ["src/**", "test/**", "pom.xml", "build.gradle*"]
  - slug: review
    description: Reviews changes made by other agents. Leaves comments only.
    tools: [read, code_search, comment]
    model: agent-review          # Cross-review with a model different from the authoring model

Rules. Organization-wide rules are written as AGENTS.md and distributed to every repository. Because most coding agents — Bob, Codex, OpenCode, and others — read this file, the rules persist even if you switch harnesses.

Subagents. Planning, implementation, and review are performed in separate contexts. The review agent should see only the diff and the plan document, not the implementation agent's conversation history, so it does not repeat the same mistakes.

Approval policies. Rather than scattering action decisions across harness settings, consolidate them into a single OPA policy. The harness (or a harness wrapper) queries the policy engine before executing a tool, and the decision is recorded in the audit log.

# platform-policy/approval.rego
package agent.approval
 
import rego.v1
 
# Decision precedence: deny > auto-allow > ask for approval
decision := "deny" if {
	denied
} else := "allow" if {
	auto_allowed
} else := "ask"
 
denied if {
	input.tool == "execute"
	some prefix in data.policy.denied_commands      # e.g., "rm -rf", "curl", "ssh"
	startswith(input.command, prefix)
}
 
denied if {
	input.tool == "edit"
	not path_allowed
}
 
auto_allowed if input.tool in {"read", "code_search"}
 
auto_allowed if {
	input.tool == "execute"
	some prefix in data.policy.approved_commands    # e.g., "mvn -q test", "git diff"
	startswith(input.command, prefix)
}
 
path_allowed if {
	some pattern in data.modes[input.mode].edit_paths
	glob.match(pattern, ["/"], input.path)
}

The advantage of this structure is that what Bob's allowed_permissions, approvedCommands, deniedCommands, and fileRegex did is managed in one place, independent of the harness, and policy changes can be reviewed as PRs.

5.3 Execution sandbox: a disposable isolated environment per task

The biggest weakness of Bob identified in Part 1 was "no system-level sandbox." In the open-source design, we put this at the center of the platform.

Loading diagram…

The key design points are as follows.

  • A new Pod per task: No files, caches, or credentials from previous tasks remain. If a build cache is needed, attach a separate read-only cache volume.
  • Kernel isolation: Specify gVisor or Kata Containers as the runtime class so that commands run by the agent do not touch the host kernel directly.
  • Least-privilege credentials: Inject a short-lived token that can push only to the work branch of the repository in question. Branch protection on main prevents direct merges.
  • Block external traffic: Deny everything by default and allow only the gateway, the Git server, and the package mirror.
# Egress policy for the sandbox namespace
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
  name: agent-sandbox-egress
  namespace: agent-sandbox
spec:
  podSelector: {}
  policyTypes: ["Egress"]
  egress:
    - to:
        - namespaceSelector:
            matchLabels: { kubernetes.io/metadata.name: model-gateway }
      ports: [{ port: 4000, protocol: TCP }]
    - to:
        - namespaceSelector:
            matchLabels: { kubernetes.io/metadata.name: git }
      ports: [{ port: 443, protocol: TCP }]
    - to:
        - namespaceSelector:
            matchLabels: { kubernetes.io/metadata.name: artifact-mirror }
      ports: [{ port: 443, protocol: TCP }]
    - to:
        - namespaceSelector:
            matchLabels: { kubernetes.io/metadata.name: kube-system }
      ports: [{ port: 53, protocol: UDP }, { port: 53, protocol: TCP }]

With this in place, even if a prompt injection leads the agent to follow an instruction like "send this file to an external address," it fails at the network layer. It is the environment that blocks it, not the agent's own judgment.

5.4 Model layer: role aliases behind a gateway

The harness knows only role aliases; the actual models are determined in the gateway configuration.

# Example LiteLLM gateway configuration (config.yaml)
model_list:
  - model_name: agent-plan            # Planning and design: prioritize reasoning performance
    litellm_params:
      model: hosted_vllm/<plan-model>
      api_base: http://vllm-plan.model-serving:8000/v1
  - model_name: agent-code            # Implementation: prioritize tool-calling quality and speed
    litellm_params:
      model: hosted_vllm/<code-model>
      api_base: http://vllm-code.model-serving:8000/v1
  - model_name: agent-review          # Review: a model different from the implementation model
    litellm_params:
      model: hosted_vllm/<review-model>
      api_base: http://vllm-review.model-serving:8000/v1
 
router_settings:
  fallbacks:
    - agent-code: ["agent-plan"]
 
litellm_settings:
  callbacks: ["otel"]
 
general_settings:
  master_key: os.environ/LITELLM_MASTER_KEY
  database_url: os.environ/DATABASE_URL   # Records virtual keys and per-team usage
  • Models by role: This is the same idea as Bob using different models for security checks, planning, and code generation. For review, assign a different model from the one used for implementation to avoid shared blind spots.
  • Per-team virtual keys and budgets: Issue keys per team and project and set usage caps. Usage is collected in the gateway database and the observability layer.
  • Model swaps: Add a new model to the gateway as an alias, send it a portion of traffic for comparison, then switch over. Harness and developer settings do not change.

Choose models using an internal evaluation set. Public benchmark scores do not guarantee performance on your own codebase, build system, and coding conventions. We recommend building an evaluation set as follows.

  1. Pick past bug-fix commits from internal repositories that include tests.
  2. Use the pre-fix state and the issue description as input, and passing the corresponding tests as the success criterion.
  3. For each "model × harness" combination, measure pass rate, tokens and time per task, number of policy blocks, and review rejection rate.

This evaluation set becomes a regression test you rerun every time you swap models or harnesses.

5.5 Code understanding layer: exposing repositories as MCP servers

Most of the time, agents get lost in large repositories because "they don't know where to look." Build code understanding capabilities as MCP servers so every harness uses them the same way.

ToolImplementationPurpose
repo_mapExtract per-file class and function signatures with tree-sitterGrasp repository structure with few tokens
find_referencesLanguage server (LSP) queriesPrecise definition and reference tracking
semantic_searchEmbedding index of code and documents (pgvector or Qdrant)Find relevant code with natural language
build_infoBuild tool metadata (modules, dependencies)Impact analysis

The index must mirror repository access permissions exactly. If code from a repository a developer cannot access leaks into search results, that in itself is a data leak.

5.6 Modernization pipeline: deterministic transformation, then the agent

We rebuild with open source the "build → analyze failures → group by root cause → fix incrementally" flow demonstrated by Bob's Java modernization mode. The key is principle 4: deterministic tools first.

Loading diagram…
  1. Analysis: Use the Konveyor analyzer to find problem areas when moving to a target technology (e.g., the latest Java, a container environment). Konveyor provides thousands of migration rules, and Kai assists with fixes by supplying these analysis results and past change history to the LLM.
  2. Deterministic transformation: Anything that can be changed mechanically is transformed in bulk with OpenRewrite recipes. The same input always produces the same result, with no model cost.
# Example: run the Java 21 migration recipe (in air-gapped environments, serve recipe artifacts from the internal Maven mirror)
mvn -U org.openrewrite.maven:rewrite-maven-plugin:run \
  -Drewrite.recipeArtifactCoordinates=org.openrewrite.recipe:rewrite-migrate-java:RELEASE \
  -Drewrite.activeRecipes=org.openrewrite.java.migrate.UpgradeToJava21
  1. Build and test to surface the remaining problems.
  2. Group by root cause: Group compile errors and test failures by error type and package. Turning dozens of errors with the same cause into a single task greatly reduces the agent's context and cost.
  3. Agent fixes: Handle one group as one sandbox task, then rebuild.
  4. Every change is left as a PR and goes through security checks and human review.

Some OpenRewrite recipe modules are provided under the Moderne Source Available License. Applying them to your organization's internal code is permitted, but they cannot be included in commercial products for resale, so you should check the license of each module you plan to use in advance.

5.7 Check and review gates

Under no circumstances does an agent's output go directly into main. Every change goes through a PR and must pass the following in CI.

CheckToolPurpose
Static security analysisSemgrep CEVulnerable code patterns (the Bob preview also used a Semgrep integration)
Secret detectiongitleaksTokens and keys introduced by the agent
Dependency and image vulnerabilitiesTrivyLibraries brought in by upgrades
TestsProject testsFunctional regressions
Cross-reviewreview mode agentReview comments from a model different from the authoring model
Human approvalBranch protection rulesFinal accountability

5.8 Governance and observability: building Bobalytics with open source

The productivity, quality, usage, and cost metrics that Bobalytics provides are assembled from three data sources.

  • Harness traces: Collect step-by-step execution per task (model calls, tool executions, policy decisions) with OpenTelemetry and view it in Langfuse. If the harness does not support OTel, supplement with gateway logs and policy engine logs.
  • Gateway usage: Tokens and cost by team, key, and model
  • Git and CI results: PRs opened, merged, and rejected; check failures

The metrics to watch on the dashboard are as follows.

MetricMeaning
Cost per merged PRToken cost ÷ number of changes actually merged — the real unit cost of productivity
PR merge rate and rejection reasonsQuality of agent output
Number of policy blocks and approval requestsWhether the policy is too loose or too strict
Approval wait timeWhether humans are the bottleneck
Success rate by modelBasis for adjusting routing

Audit logs — policy decisions and approval records — are kept separately in immutable storage (e.g., object storage configured for WORM). Observability data is for analysis, while audit logs are for accountability, so their retention policies differ.

5.9 Air-gapped supply chain: the import procedure is the operating capability

The hidden core of an air-gapped platform is what you bring in from outside, and through what procedure.

What is importedInternal repositoryExample update cadence
Container images (harness, gateway, serving, tools)HarborMonthly + security patches as needed
Language packages (Maven, npm, PyPI)Package mirrorOn project request
Model weightsInternal object storageWhen evaluation is passed
Editor extensions (.vsix)Internal distribution repositoryQuarterly
Vulnerability databases and security rules (Trivy DB, Semgrep rules)Internal mirrorWeekly
OpenRewrite recipe artifactsMaven mirrorDuring modernization work

Standardize the import procedure in this order: "download in the external staging zone → verify signatures and hashes → scan for vulnerabilities → approve → import internally → run regression tests on the internal evaluation set → deploy." Model weights and the harness in particular change the agent's behavior when they change, so deploy only combinations that have passed regression testing.

6. Comparing Bob with the open-source approach

AspectIBM Bob (self-hosted)Open-source approach
Speed of adoptionFast — an integrated productSlow — requires assembly and integration
IsolationCentered on mechanisms inside the agent; outer isolation must be designed separatelyDisposable sandbox at the center of the design
Model choiceCentered on supported models (two for air-gapped at general availability)Any open-weight model the license permits
Policy managementBob's configuration formatA single OPA policy, managed independently of the harness
ModernizationJava, IBM i, and Z premium packagesJava can be assembled with OpenRewrite and Konveyor; few alternatives for IBM i and Z
ObservabilityBobalyticsAssembled from OTel, Langfuse, and Grafana
Support and accountabilityIBMInternal team (or a support contract)
Continuity riskVendor policy changesIndividual projects being discontinued (e.g., Roo Code) — mitigated by standards-based layer separation

Rather than one being the right answer, it is a trade-off between control and operational burden. If you have limited operations staff and IBM platform assets are core, Bob fits; if you have platform engineering capability and need to control models, isolation, and cost directly, the open-source approach fits. Mixing the two is also possible. For example, even if you use Bob, the sandbox (5.3), gateway (5.4), and check gates (5.7) from this post can be applied as-is as outer layers.

7. Phased roadmap

PhaseScopeCompletion criteria
0. EvaluationBuild the internal evaluation set; compare 2–3 harnesses × candidate modelsStandard harness and per-role model candidates finalized
1. MVPGateway + model serving + standard harness (CLI) + disposable sandbox + PR check gatesPilot team merges real work via PRs; external traffic blocking verified
2. ExpansionACP editor integration, browser IDE, modes and OPA policies as code, observability dashboardPer-team budgets and policies in operation; cost per merged PR measured
3. MaturityCode understanding MCP servers, Java modernization pipeline, multi-team onboarding, automated import procedureOne modernization project completed; quarterly model swaps carried out via regression testing

Some things were deliberately left out of the MVP. Editor integration and the modernization pipeline are highly valuable, but expanding them first without isolation and check gates makes it hard to roll back. We recommend putting safeguards in place first and then broadening the scope of use.

8. Risks and mitigations

RiskMitigation
Open-weight model quality falls short of commercial modelsRole-based routing, deterministic tools first, workflows that break tasks into small pieces, managing expectations with an internal evaluation set
Model license violationsInclude a license review step in the import procedure; record permitted uses per model
Open-source projects being discontinuedSeparate layers with standard protocols; continuously validate harness swaps with the evaluation set
Data leakage or destruction via prompt injectionBlock external traffic, disposable sandboxes, least-privilege credentials, no direct merges to main
Policies so strict that nobody uses the platformWatch approval-request and block metrics and relax policies incrementally
Lack of operations staffLimit scope by phase, automate the import procedure, external support contracts if needed

Closing thoughts

What IBM Bob demonstrates is that the competitive edge of development agents is shifting from the model to the platform. That platform consists of the harness, modes and policies, isolation, model routing, auditing, modernization workflows, and the air-gapped supply chain.

Most of these components can be built with open source. However, the success of an open-source approach depends less on the choice of individual tools than on the design principles of separating layers with standards, delegating isolation to the environment, and managing policies and evaluation as code. Whether you adopt Bob or build your own, these principles apply just the same.

References