Adapts modernization paper
Evidence-backed mainframe modernization at enterprise scale.
How a large bank validated deterministic system discovery across a complex mainframe estate — separating system discovery from AI reasoning before expanding modernization decisions.
Based on an anonymized customer engagement. The customer is not named. Metrics reflect that project’s outcomes; results vary by estate. Quantitative claims are limited to the measured 100-program validation described below.
01 · Situation
Useful on small files. Unreliable at estate scale.
In 2026, a large bank began evaluating approaches for modernizing an estate of 14,155 mission-critical mainframe applications spanning COBOL, HLASM, PL/I, IDMS, CICS, BMS, ADSO, DB2, Java, and cross-application dependencies. General-purpose AI assistants were useful when engineers worked with individual files or isolated repositories. Reliability became harder to maintain when questions crossed programs, languages, databases, transactions, and application boundaries.
Why IDMS became a critical test
The IDMS stress test.
IDMS exposed the limitations of relying on AI interpretation alone. The environment combined COBOL and Assembler-related constructs with IDMS-specific syntax, dialogs, database relationships, and runtime behavior. CICS macros, BMS maps, ADSO dialogs, DB2 calls, jobs, and other mainframe artifacts added further context.
Client environment
A connected estate, not a collection of source files.
The bank operated a large and heterogeneous mainframe environment. Application behavior rarely existed within a single program.
- COBOL
- HLASM
- PL/I
- IDMS
- CICS
- BMS
- ADSO
- DB2
- Java
- Shared libraries and copybooks
- Jobs and started tasks
- Cross-repository and cross-application dependencies
A transaction could begin through a CICS interface, traverse COBOL or PL/I programs, interact with IDMS or DB2, invoke shared components, participate in batch processes, and affect applications owned by different teams. Modernization therefore required understanding the system as a connected estate.
02 · Challenge
Code understanding is not system understanding.
The core challenge was not whether AI could explain COBOL or summarize a program. The challenge was whether an engineering team could reliably answer questions that span the estate.
Core questions
- How does a business transaction flow across multiple programs?
- Which IDMS entities and data structures participate in the flow?
- Where are business rules implemented?
- What changes when a field, program, interface, or dependency changes?
- Why does a technical capability exist, and what business process does it support?
Consequential questions
- What calls the program?
- What does the program call?
- Which data structures does it read or modify?
- Which transactions or jobs depend on it?
- Where else is the same business rule implemented?
- Which downstream systems could be affected by a change?
- What business capability justifies the program’s continued existence?
For a banking environment, plausible answers were insufficient. Engineering teams required answers that could be verified against the underlying system.
Where trust breaks
From mainframe estate to a trust and risk problem.
- Mainframe estate
COBOL, HLASM, PL/I, IDMS, CICS, BMS, ADSO, DB2, Java
- Generic AI
Works on isolated code samples
- Cross-program analysis needed
Transaction flows, shared dependencies
- Hallucinations
Outputs not traceable to code
- Trust & risk problem
Must be accurate, auditable
03 · Validation design
Map first. Reason second.
The bank selected 100 programs for a focused validation. Five participants evaluated Adapts using approximately 50–100 questions designed to test different forms of system understanding.
Establish what the system contains and how it is connected before asking AI to explain or transform it.
| Validation area | Bank question | Grounding needed |
|---|---|---|
| Program flow understanding | Can Adapts reconstruct how execution moves across programs and related components — calls, transaction paths, jobs, shared components, upstream and downstream flows? | Explain a workflow rather than merely summarize an individual program. |
| IDMS and data relationships | Can Adapts identify relationships between programs and the underlying data environment — IDMS structures, database relationships, fields and records, shared data dependencies? | Data relationships frequently connect programs that appear independent when viewed only through source-code boundaries. |
| Business rule identification | Can Adapts identify business rules embedded in the implementation — validations, calculations, decision logic, eligibility, exceptions — and where else those rules appear? | For modernization, identifying a rule is only part of the problem. Teams also need to know which processes depend on it. |
| Change and impact analysis | Can a proposed change to a field, program, data structure, interface, or shared dependency be followed through the dependency graph? | Identify the blast radius across programs, data stores, screens, transactions, and other dependent components. |
| Business context understanding | Can technical behavior be related to the business capability or process it supports — moving from “what does the code do?” to “why does the behavior exist?” | Technical existence alone does not establish whether a capability remains necessary during modernization. |
Inside Adapts for mainframe estates
Three engines, one grounded mainframe context
The same three-engine architecture, re-grounded for the estate. Static parsers turn IDMS, COBOL, PL/I, DB2, HLASM, and JCL into one deterministic graph — from a single program to the whole enterprise.
Holds the parsed mainframe estate — not natural-language guesses — so answers are grounded in real IDMS, COBOL, PL/I, DB2, HLASM, and JCL.
- Static parsers: IDMS, COBOL, PL/I, DB2, HLASM, JCL
- Copybooks, PDS members & CICS/BMS/ADSO artifacts
- Latest per load module & compile
Stores the deterministic graph of how everything connects — from a single program or copybook up to jobs, transactions, and the whole enterprise.
- Deterministic connection graph
- Program/copybook → application → enterprise scale
- Job, started task & CICS transaction linkage
Enriches the picture with context beyond the source: runbooks, tickets, job schedulers, and mainframe runtime signals.
- Runbooks, tickets & specs
- Job schedulers & SMF/CICS runtime signals
- Third-party systems e.g. Jira, Confluence, Notion, ADO, etc.
- Onboarding
- Architectural Diagrams
- Drift Analysis
- Gap Analysis
- Code Review and Analysis
- Implementation Plans
- Modernization Plans
- Scrum Boards
- Implementation Suites
- Test Suites
- Infrastructure as Code
- DevOps
- Release Notes
04 · Deterministic approach
Build solid ground before the model speaks.
Stage 1
Discover the technical estate
Process available source artifacts and extract identifiable technical elements before AI reasoning begins — programs, functions, variables, interfaces, copybooks, database references, jobs, transactions, maps, configuration, and shared libraries. The inventory provides the baseline for understanding what exists.
Stage 2
Build the relationship graph
Connect discovered elements into a persistent graph — program-to-program calls, copybook relationships, database interactions, jobs, CICS transactions, shared-library usage, and cross-application dependencies. Relationships can be traversed rather than rediscovered through every AI prompt.
Stage 3
Enrich the technical model
Layer in evidence beyond source: specifications, runbooks, tickets, job schedulers, runtime signals, and systems such as Jira, Confluence, and Azure DevOps. Such evidence helps connect technical behavior with operational and business context.
Stage 4
Apply AI reasoning over grounded context
AI reasons after the technical model is established. Instead of asking the model to infer relationships independently, Adapts provides relevant programs, dependencies, data structures, business context, and supporting evidence. The model reasons over an already discovered enterprise model.
| Dimension | LLM-only risk | Adapts grounding |
|---|---|---|
| Discovery method | Prompt-driven inference over partial files or repository context. | Static analysis first; graph construction before LLM reasoning. |
| Context boundary | Partial files and repo windows that break across systems. | Connected system map spanning repositories, languages, and dependencies. |
| IDMS handling | Risk of treating IDMS syntax as natural-language or COBOL-like context. | Deterministic parsing/extraction for IDMS constructs, then LLM explanation on verified facts. |
| Output quality risk | Plausible explanations may not be traceable to source relationships. | Insights linked to programs, variables, calls, maps, and graph edges. |
| User trust model | “Does the answer sound correct?” | “What evidence supports the answer?” |
The pattern
What fails. What works.
LLM-only
- Prompt-driven inference
Partial files, repo context
- Context degrades
Fails to scale across systems
- Hallucination risk
Explanations not traceable
Adapts
- Static analysis first
Graph built before LLM
- Connected system map
Repos, languages, dependencies
- Evidence-based reasoning
>95% fewer hallucinations
05 · Validation results
Reviewer-rated accuracy of 91–98%.
Across the 100-program evaluation, five reviewers assessed approximately 50–100 questions spanning the five validation areas. Reviewer-rated accuracy ranged from 91% to 98%.
The range reflects reviewer evaluation of answers against their knowledge of the system and available technical evidence. The result is particularly relevant because the evaluation extended beyond isolated source-code explanation — covering program flows, IDMS and data relationships, business rules, cross-program dependencies, change impact, and business context.
The validation measured whether Adapts could maintain useful accuracy while moving from program understanding toward system understanding.
What the validation demonstrated
From program understanding toward system understanding.
Cross-program context could be preserved
Program understanding did not stop at source-file boundaries. The graph-based model allowed questions to traverse related programs, calls, data structures, jobs, and dependencies.
IDMS could be treated as system structure rather than prompt context
IDMS-specific constructs and relationships could be incorporated into the discovered model before generative reasoning — reducing dependence on an LLM correctly interpreting unfamiliar mainframe syntax from isolated snippets.
Business rules could be connected to implementation
Rules embedded inside legacy programs could be identified and associated with the technical paths through which they were applied — important when deciding whether behavior should be preserved, modified, consolidated, or retired.
Change analysis could use actual dependency relationships
The model could follow upstream and downstream relationships when reviewers asked about the potential effect of a change. Impact analysis became traversing known relationships rather than relying only on model inference.
Business context could be evaluated alongside technical behavior
The validation included questions about the business purpose supported by technical implementation — creating a bridge to the later modernization question: does the behavior still need to exist?
Evidence instead of plausibility
Answer quality is not the same as answer traceability.
A generated explanation may sound reasonable without being correct. Adapts treats supporting evidence as part of the answer.
- Programs
- Source locations
- Variables
- Calls
- Database relationships
- Graph edges
- Transactions
- Jobs
- Business documentation
The reviewer can inspect why a conclusion was reached rather than being required to trust the language model’s interpretation. The trust model changes from “Does the answer sound correct?” to “What evidence supports the answer?”
From discovery to modernization
Discovery should not end as static documentation.
A persistent system model can support multiple phases of modernization. Continuity matters: system knowledge created during discovery should remain available to the engineers and AI agents performing the transformation.
Understand
- System documentation
- Onboarding
- Architecture analysis
- Business-rule discovery
- Dependency analysis
- Gap analysis
- Change-impact analysis
Plan
- Modernization scope
- Implementation planning
- Transformation plans
- Work decomposition
- Engineering backlogs
Build and validate
- APIs and IDE integrations
- MCP servers and chat interfaces
- Coding agents
- Implementation and testing
- Infrastructure, DevOps, and release workflows
Case study at a glance
Measured validation across technical and business understanding.
| Measure | Result |
|---|---|
| Industry | Banking |
| Environment | Enterprise mainframe |
| Broader estate | 14,155 applications |
| Technology landscape | COBOL, HLASM, PL/I, IDMS, CICS, BMS, ADSO, DB2, Java |
| Measured validation scope | 100 programs |
| Reviewers | 5 |
| Validation questions | Approximately 50–100 |
| Question categories | Program flows, IDMS/data relationships, business rules, change analysis, business context |
| Reviewer-rated accuracy | 91–98% |
| Core architecture | Deterministic discovery and relationship mapping before AI reasoning |
| Primary outcome | Evidence-backed system understanding across programs, dependencies, data, business rules, and business context |
Scope and limitations
A measured validation — not a completed estate map.
The validation covered 100 programs, not the bank’s complete estate of 14,155 applications. Results should be interpreted as a measured validation of the approach rather than proof that every application had already been understood.
Reviewer-rated accuracy of 91–98% reflects the questions and programs included in the evaluation. The result should not be generalized to every technology, application, or modernization environment without equivalent validation.
Runtime evidence may vary by environment. Legacy systems frequently contain incomplete logging, inaccessible telemetry, or processes executed only under specific business conditions. Absence of observed activity should not be interpreted as evidence that functionality is unnecessary.
Such limitations reinforce the need for confidence scoring, evidence visibility, and targeted SME validation as system discovery expands.
06 · Core lesson
System discovery establishes the evidence. AI reasons over the evidence.
The engagement demonstrated that the principal limitation of AI-led modernization was not whether an LLM could understand code. The limitation appeared when the model was expected to reconstruct a large enterprise system from fragmented context.
For the bank, the 100-program validation provided a measurable basis for evaluating that approach: five reviewers, approximately 50–100 questions spanning technical and business understanding, and reviewer-rated accuracy between 91% and 98%.
From: Can AI understand our code?
To: Do we have enough verified system context for engineers and AI to make reliable modernization decisions?
See it on your codebase
See Adapts on a mainframe-like estate.
30-minute technical walkthrough with an enterprise architect. No slides · a live demo on real code.