Inside IBM Bob (Part 2) — Designing an Open-Source Coding Agent Platform for Air-Gapped Environments
We mapped IBM Bob's components to open-source equivalents one by one to design a development agent platform that can run in an air-gapped environment with no external connectivity. It covers layer-by-layer design — harness, sandbox, model gateway, policy, modernization pipeline, observability, and supply chain — along with configuration examples and a phased roadmap.
In Part 1, we analyzed IBM Bob. In short, Bob's value lies less in the model itself than in its common harness, role-based modes and approval policies, auditing, modernization workflows, and a deployment model that brings all of this inside the customer's air-gapped network.
This post covers how we would design the same platform if we built it ourselves with open source. The goal is not to clone Bob as-is, but to use the structure Bob demonstrates as a baseline while addressing the weaknesses identified in Part 1 (no system-level sandbox, limited model choice in air-gapped environments).
This post is a design document. The configuration examples are meant to illustrate the structure and should be validated against each project's documentation before being used in production. Licenses and project status for the open-source projects are as of October 4, 2026, and may change.
1. Requirements
We start by laying out the requirements an air-gapped development agent platform must meet.
| Area | Requirement |
|---|---|
| Data | Source code, prompts, and logs never leave the organization's boundary |
| Isolation | Commands the agent runs execute only inside an isolated environment and cannot communicate with anything other than approved internal services |
| Permissions | The tools available and the files that can be modified differ by role (ask, plan, implement, review) |
| Approval | State-changing actions are judged by policy as auto-allowed, approval-required, or blocked |
| Auditing | It is possible to reconstruct who (which agent) did what, and on what basis |
| Cost | Usage can be measured per team and project, with caps applied |
| Models | Not tied to any specific model; models can be switched by task type |
| Surfaces | Usable both in developers' editors and in CI pipelines |
| Modernization | Large-scale changes such as framework and JDK upgrades are carried out as repeatable procedures |
| Operations | Models, tools, and vulnerability databases can be updated without external connectivity |
2. Design principles
We condensed the requirements into five principles. Every layer design that follows adheres to them.
- Separate layers with standard protocols. ACP between editors and agents, MCP between agents and tools, an OpenAI-compatible API between agents and models, and
AGENTS.mdfor rules. When the boundaries between layers are standard, if an open-source project in one layer stalls, you only need to replace that layer. - Isolate outside the agent. An agent's internal ignore lists and approval settings are conveniences, not security boundaries. The security boundary is the responsibility of the container runtime and network policies.
- Hide models behind a gateway. The harness knows only role aliases such as
agent-planandagent-code, not actual model names. Model swaps, routing, budgets, and logging are handled at the gateway. - Deterministic tools first, models for the rest. Whatever can be done with compilers, tests, static analysis, and code transformation tools should be done with them; models are used to interpret those results and fix the remaining problems.
- Manage policies and configuration in Git. Mode definitions, approval policies, rules, and gateway configuration all live in a repository and are deployed after review. Policy changes are subject to auditing too.
3. Overall architecture
Compared with Bob's structure from Part 1, the biggest difference is that Bob's "common harness" is split into ② and ③, separating agent execution from the command execution environment.
4. Mapping Bob's components to open source
| Bob component | Open-source candidates | License | What to check when choosing |
|---|---|---|---|
| Harness + BobShell | OpenAI Codex CLI, OpenCode, Goose | Apache-2.0, MIT, Apache-2.0 | Non-interactive execution, custom model endpoints, MCP and ACP support, approval hooks |
| Bob IDE | code-server + Cline or Continue | MIT, Apache-2.0 | The extension's mode and approval features, air-gapped installation (bringing in extension files) |
| Custom mode format | Roo Code–family extensions | Apache-2.0 | Roo Code was archived in May 2026. Check the maintenance status of community forks |
| Editor integration | Agent Client Protocol (ACP) | Apache-2.0 | Confirm ACP support across all harness candidates |
| Tool integration | Model Context Protocol (MCP) servers | Varies by server | Authentication and permission scope, audit logs |
| Model routing | LiteLLM | MIT (some enterprise features separate) | Whether the features you need (virtual keys, budgets, SSO) fall within the open-source scope |
| Model serving | vLLM, SGLang | Apache-2.0 | Tool-call parser support, long-context performance |
| Models | Open-weight coding models (Qwen coder family, Mistral Devstral family, OpenAI gpt-oss, NVIDIA Nemotron, etc.) | Varies by model | Terms for commercial use, tool-calling quality, results on your internal evaluation set |
| Code understanding | tree-sitter, language servers (LSP), pgvector · Qdrant | MIT, varies by server, PostgreSQL · Apache-2.0 | Indexing time for large repositories, respecting access permissions |
| Java modernization package | OpenRewrite, Konveyor (including Kai) | Core Apache-2.0 (some recipe modules under the Moderne Source Available License), Apache-2.0 | Per-module recipe licenses — some modules can be applied internally but cannot be resold in commercial products |
| Inline security checks | Semgrep CE, gitleaks, Trivy | LGPL-2.1, MIT, Apache-2.0 | Offline updates for rules and vulnerability databases |
| Sandbox | Kubernetes/OpenShift + gVisor or Kata Containers | Apache-2.0 | Runtime compatibility, build tool performance |
| Approval policy | Open Policy Agent (OPA) | Apache-2.0 | Integration point with the harness |
| Bobalytics | OpenTelemetry, Langfuse, Grafana | Apache-2.0, MIT (core), AGPL-3.0 | Grafana is AGPL — internal use is generally fine, but check your policy |
| Air-gapped deployment | Harbor (images), package mirrors (Nexus Repository, etc.) | Apache-2.0, varies by product | Usage terms for the mirror product's free edition |
There is effectively no mature open-source equivalent to the IBM i (RPG) and mainframe (Z) packages. We explicitly leave this area as a limitation of the open-source approach.
5. Layer-by-layer design
5.1 User surfaces: ACP for editors, non-interactive for CI
There are three user surfaces.
- Editors: Editors that support ACP (Zed, JetBrains, Neovim, etc.) connect by launching the harness CLI as an ACP server. The editor renders only the UI and approval dialogs, while the harness handles execution. This is the same structure Bob uses with
bob acp. - Browser IDE: In air-gapped environments where installing tools on developer PCs is difficult, run code-server internally and distribute an image with the agent extension preinstalled. Extension files (
.vsix) are brought in through the supply chain procedure (5.9). - CI: Repetitive tasks (dependency upgrades, documentation updates, draft reviews) run in pipelines using the harness's non-interactive mode, and results are always left as PRs.
Standardize on one harness, but keep it replaceable. The comparison criteria are as follows.
| Criterion | How to check |
|---|---|
| Custom model endpoints | Can you configure the OpenAI-compatible gateway address and role aliases? |
| Approval hooks | Can it call an external policy engine before executing a tool? (If not, a wrapper is needed) |
| Rules file | Does it read AGENTS.md? |
| Non-interactive execution | Does it reliably produce exit codes and artifacts in CI? |
| Observability | Does it export trace data via OpenTelemetry? |
| Project health | Release cadence, number of contributors, governance (foundation membership) |
The last criterion has a real-world precedent. Roo Code, which used a format nearly identical to Bob's mode format, had its repository archived by the original development team in May 2026. This is why principle 1 — separating layers with standards — matters.
5.2 Harness: modes, rules, and approvals as code
Modes. Since every harness has a different mode configuration format, the platform keeps its source definitions in a neutral format in Git and converts them into each harness's format for deployment. Here is an example source definition modeled on Bob's format.
# platform-policy/modes.yaml — platform standard modes (source definition)
modes:
- slug: ask
description: Code analysis and Q&A. Changes nothing.
tools: [read, code_search]
- slug: plan
description: Change planning. Writes plan documents only.
tools: [read, code_search, edit]
edit_paths: ["docs/plans/**"]
- slug: code
description: Implementation and testing.
tools: [read, code_search, edit, execute, mcp]
edit_paths: ["src/**", "test/**", "pom.xml", "build.gradle*"]
- slug: review
description: Reviews changes made by other agents. Leaves comments only.
tools: [read, code_search, comment]
model: agent-review # Cross-review with a model different from the authoring modelRules. Organization-wide rules are written as AGENTS.md and distributed to every repository. Because most coding agents — Bob, Codex, OpenCode, and others — read this file, the rules persist even if you switch harnesses.
Subagents. Planning, implementation, and review are performed in separate contexts. The review agent should see only the diff and the plan document, not the implementation agent's conversation history, so it does not repeat the same mistakes.
Approval policies. Rather than scattering action decisions across harness settings, consolidate them into a single OPA policy. The harness (or a harness wrapper) queries the policy engine before executing a tool, and the decision is recorded in the audit log.
# platform-policy/approval.rego
package agent.approval
import rego.v1
# Decision precedence: deny > auto-allow > ask for approval
decision := "deny" if {
denied
} else := "allow" if {
auto_allowed
} else := "ask"
denied if {
input.tool == "execute"
some prefix in data.policy.denied_commands # e.g., "rm -rf", "curl", "ssh"
startswith(input.command, prefix)
}
denied if {
input.tool == "edit"
not path_allowed
}
auto_allowed if input.tool in {"read", "code_search"}
auto_allowed if {
input.tool == "execute"
some prefix in data.policy.approved_commands # e.g., "mvn -q test", "git diff"
startswith(input.command, prefix)
}
path_allowed if {
some pattern in data.modes[input.mode].edit_paths
glob.match(pattern, ["/"], input.path)
}The advantage of this structure is that what Bob's allowed_permissions, approvedCommands, deniedCommands, and fileRegex did is managed in one place, independent of the harness, and policy changes can be reviewed as PRs.
5.3 Execution sandbox: a disposable isolated environment per task
The biggest weakness of Bob identified in Part 1 was "no system-level sandbox." In the open-source design, we put this at the center of the platform.
The key design points are as follows.
- A new Pod per task: No files, caches, or credentials from previous tasks remain. If a build cache is needed, attach a separate read-only cache volume.
- Kernel isolation: Specify gVisor or Kata Containers as the runtime class so that commands run by the agent do not touch the host kernel directly.
- Least-privilege credentials: Inject a short-lived token that can push only to the work branch of the repository in question. Branch protection on main prevents direct merges.
- Block external traffic: Deny everything by default and allow only the gateway, the Git server, and the package mirror.
# Egress policy for the sandbox namespace
apiVersion: networking.k8s.io/v1
kind: NetworkPolicy
metadata:
name: agent-sandbox-egress
namespace: agent-sandbox
spec:
podSelector: {}
policyTypes: ["Egress"]
egress:
- to:
- namespaceSelector:
matchLabels: { kubernetes.io/metadata.name: model-gateway }
ports: [{ port: 4000, protocol: TCP }]
- to:
- namespaceSelector:
matchLabels: { kubernetes.io/metadata.name: git }
ports: [{ port: 443, protocol: TCP }]
- to:
- namespaceSelector:
matchLabels: { kubernetes.io/metadata.name: artifact-mirror }
ports: [{ port: 443, protocol: TCP }]
- to:
- namespaceSelector:
matchLabels: { kubernetes.io/metadata.name: kube-system }
ports: [{ port: 53, protocol: UDP }, { port: 53, protocol: TCP }]With this in place, even if a prompt injection leads the agent to follow an instruction like "send this file to an external address," it fails at the network layer. It is the environment that blocks it, not the agent's own judgment.
5.4 Model layer: role aliases behind a gateway
The harness knows only role aliases; the actual models are determined in the gateway configuration.
# Example LiteLLM gateway configuration (config.yaml)
model_list:
- model_name: agent-plan # Planning and design: prioritize reasoning performance
litellm_params:
model: hosted_vllm/<plan-model>
api_base: http://vllm-plan.model-serving:8000/v1
- model_name: agent-code # Implementation: prioritize tool-calling quality and speed
litellm_params:
model: hosted_vllm/<code-model>
api_base: http://vllm-code.model-serving:8000/v1
- model_name: agent-review # Review: a model different from the implementation model
litellm_params:
model: hosted_vllm/<review-model>
api_base: http://vllm-review.model-serving:8000/v1
router_settings:
fallbacks:
- agent-code: ["agent-plan"]
litellm_settings:
callbacks: ["otel"]
general_settings:
master_key: os.environ/LITELLM_MASTER_KEY
database_url: os.environ/DATABASE_URL # Records virtual keys and per-team usage- Models by role: This is the same idea as Bob using different models for security checks, planning, and code generation. For review, assign a different model from the one used for implementation to avoid shared blind spots.
- Per-team virtual keys and budgets: Issue keys per team and project and set usage caps. Usage is collected in the gateway database and the observability layer.
- Model swaps: Add a new model to the gateway as an alias, send it a portion of traffic for comparison, then switch over. Harness and developer settings do not change.
Choose models using an internal evaluation set. Public benchmark scores do not guarantee performance on your own codebase, build system, and coding conventions. We recommend building an evaluation set as follows.
- Pick past bug-fix commits from internal repositories that include tests.
- Use the pre-fix state and the issue description as input, and passing the corresponding tests as the success criterion.
- For each "model × harness" combination, measure pass rate, tokens and time per task, number of policy blocks, and review rejection rate.
This evaluation set becomes a regression test you rerun every time you swap models or harnesses.
5.5 Code understanding layer: exposing repositories as MCP servers
Most of the time, agents get lost in large repositories because "they don't know where to look." Build code understanding capabilities as MCP servers so every harness uses them the same way.
| Tool | Implementation | Purpose |
|---|---|---|
repo_map | Extract per-file class and function signatures with tree-sitter | Grasp repository structure with few tokens |
find_references | Language server (LSP) queries | Precise definition and reference tracking |
semantic_search | Embedding index of code and documents (pgvector or Qdrant) | Find relevant code with natural language |
build_info | Build tool metadata (modules, dependencies) | Impact analysis |
The index must mirror repository access permissions exactly. If code from a repository a developer cannot access leaks into search results, that in itself is a data leak.
5.6 Modernization pipeline: deterministic transformation, then the agent
We rebuild with open source the "build → analyze failures → group by root cause → fix incrementally" flow demonstrated by Bob's Java modernization mode. The key is principle 4: deterministic tools first.
- Analysis: Use the Konveyor analyzer to find problem areas when moving to a target technology (e.g., the latest Java, a container environment). Konveyor provides thousands of migration rules, and Kai assists with fixes by supplying these analysis results and past change history to the LLM.
- Deterministic transformation: Anything that can be changed mechanically is transformed in bulk with OpenRewrite recipes. The same input always produces the same result, with no model cost.
# Example: run the Java 21 migration recipe (in air-gapped environments, serve recipe artifacts from the internal Maven mirror)
mvn -U org.openrewrite.maven:rewrite-maven-plugin:run \
-Drewrite.recipeArtifactCoordinates=org.openrewrite.recipe:rewrite-migrate-java:RELEASE \
-Drewrite.activeRecipes=org.openrewrite.java.migrate.UpgradeToJava21- Build and test to surface the remaining problems.
- Group by root cause: Group compile errors and test failures by error type and package. Turning dozens of errors with the same cause into a single task greatly reduces the agent's context and cost.
- Agent fixes: Handle one group as one sandbox task, then rebuild.
- Every change is left as a PR and goes through security checks and human review.
Some OpenRewrite recipe modules are provided under the Moderne Source Available License. Applying them to your organization's internal code is permitted, but they cannot be included in commercial products for resale, so you should check the license of each module you plan to use in advance.
5.7 Check and review gates
Under no circumstances does an agent's output go directly into main. Every change goes through a PR and must pass the following in CI.
| Check | Tool | Purpose |
|---|---|---|
| Static security analysis | Semgrep CE | Vulnerable code patterns (the Bob preview also used a Semgrep integration) |
| Secret detection | gitleaks | Tokens and keys introduced by the agent |
| Dependency and image vulnerabilities | Trivy | Libraries brought in by upgrades |
| Tests | Project tests | Functional regressions |
| Cross-review | review mode agent | Review comments from a model different from the authoring model |
| Human approval | Branch protection rules | Final accountability |
5.8 Governance and observability: building Bobalytics with open source
The productivity, quality, usage, and cost metrics that Bobalytics provides are assembled from three data sources.
- Harness traces: Collect step-by-step execution per task (model calls, tool executions, policy decisions) with OpenTelemetry and view it in Langfuse. If the harness does not support OTel, supplement with gateway logs and policy engine logs.
- Gateway usage: Tokens and cost by team, key, and model
- Git and CI results: PRs opened, merged, and rejected; check failures
The metrics to watch on the dashboard are as follows.
| Metric | Meaning |
|---|---|
| Cost per merged PR | Token cost ÷ number of changes actually merged — the real unit cost of productivity |
| PR merge rate and rejection reasons | Quality of agent output |
| Number of policy blocks and approval requests | Whether the policy is too loose or too strict |
| Approval wait time | Whether humans are the bottleneck |
| Success rate by model | Basis for adjusting routing |
Audit logs — policy decisions and approval records — are kept separately in immutable storage (e.g., object storage configured for WORM). Observability data is for analysis, while audit logs are for accountability, so their retention policies differ.
5.9 Air-gapped supply chain: the import procedure is the operating capability
The hidden core of an air-gapped platform is what you bring in from outside, and through what procedure.
| What is imported | Internal repository | Example update cadence |
|---|---|---|
| Container images (harness, gateway, serving, tools) | Harbor | Monthly + security patches as needed |
| Language packages (Maven, npm, PyPI) | Package mirror | On project request |
| Model weights | Internal object storage | When evaluation is passed |
Editor extensions (.vsix) | Internal distribution repository | Quarterly |
| Vulnerability databases and security rules (Trivy DB, Semgrep rules) | Internal mirror | Weekly |
| OpenRewrite recipe artifacts | Maven mirror | During modernization work |
Standardize the import procedure in this order: "download in the external staging zone → verify signatures and hashes → scan for vulnerabilities → approve → import internally → run regression tests on the internal evaluation set → deploy." Model weights and the harness in particular change the agent's behavior when they change, so deploy only combinations that have passed regression testing.
6. Comparing Bob with the open-source approach
| Aspect | IBM Bob (self-hosted) | Open-source approach |
|---|---|---|
| Speed of adoption | Fast — an integrated product | Slow — requires assembly and integration |
| Isolation | Centered on mechanisms inside the agent; outer isolation must be designed separately | Disposable sandbox at the center of the design |
| Model choice | Centered on supported models (two for air-gapped at general availability) | Any open-weight model the license permits |
| Policy management | Bob's configuration format | A single OPA policy, managed independently of the harness |
| Modernization | Java, IBM i, and Z premium packages | Java can be assembled with OpenRewrite and Konveyor; few alternatives for IBM i and Z |
| Observability | Bobalytics | Assembled from OTel, Langfuse, and Grafana |
| Support and accountability | IBM | Internal team (or a support contract) |
| Continuity risk | Vendor policy changes | Individual projects being discontinued (e.g., Roo Code) — mitigated by standards-based layer separation |
Rather than one being the right answer, it is a trade-off between control and operational burden. If you have limited operations staff and IBM platform assets are core, Bob fits; if you have platform engineering capability and need to control models, isolation, and cost directly, the open-source approach fits. Mixing the two is also possible. For example, even if you use Bob, the sandbox (5.3), gateway (5.4), and check gates (5.7) from this post can be applied as-is as outer layers.
7. Phased roadmap
| Phase | Scope | Completion criteria |
|---|---|---|
| 0. Evaluation | Build the internal evaluation set; compare 2–3 harnesses × candidate models | Standard harness and per-role model candidates finalized |
| 1. MVP | Gateway + model serving + standard harness (CLI) + disposable sandbox + PR check gates | Pilot team merges real work via PRs; external traffic blocking verified |
| 2. Expansion | ACP editor integration, browser IDE, modes and OPA policies as code, observability dashboard | Per-team budgets and policies in operation; cost per merged PR measured |
| 3. Maturity | Code understanding MCP servers, Java modernization pipeline, multi-team onboarding, automated import procedure | One modernization project completed; quarterly model swaps carried out via regression testing |
Some things were deliberately left out of the MVP. Editor integration and the modernization pipeline are highly valuable, but expanding them first without isolation and check gates makes it hard to roll back. We recommend putting safeguards in place first and then broadening the scope of use.
8. Risks and mitigations
| Risk | Mitigation |
|---|---|
| Open-weight model quality falls short of commercial models | Role-based routing, deterministic tools first, workflows that break tasks into small pieces, managing expectations with an internal evaluation set |
| Model license violations | Include a license review step in the import procedure; record permitted uses per model |
| Open-source projects being discontinued | Separate layers with standard protocols; continuously validate harness swaps with the evaluation set |
| Data leakage or destruction via prompt injection | Block external traffic, disposable sandboxes, least-privilege credentials, no direct merges to main |
| Policies so strict that nobody uses the platform | Watch approval-request and block metrics and relax policies incrementally |
| Lack of operations staff | Limit scope by phase, automate the import procedure, external support contracts if needed |
Closing thoughts
What IBM Bob demonstrates is that the competitive edge of development agents is shifting from the model to the platform. That platform consists of the harness, modes and policies, isolation, model routing, auditing, modernization workflows, and the air-gapped supply chain.
Most of these components can be built with open source. However, the success of an open-source approach depends less on the choice of individual tools than on the design principles of separating layers with standards, delegating isolation to the environment, and managing policies and evaluation as code. Whether you adopt Bob or build your own, these principles apply just the same.
References
- Part 1 — Inside IBM Bob: What Makes an Agentic Development Platform Different
- Agent Client Protocol · Model Context Protocol · AGENTS.md
- LiteLLM · vLLM · SGLang
- OpenAI Codex CLI · OpenCode · Goose · Cline · Continue · code-server
- OpenRewrite · OpenRewrite licensing · Konveyor AI
- Open Policy Agent · gVisor · Kata Containers
- Semgrep · gitleaks · Trivy
- OpenTelemetry · Langfuse · Harbor