Skip to content

About

Agentic first-pass review of SAP EarlyWatch Alert reports: a Python/LangGraph CLI that turns supported Word XML reports into Excel workbooks using Azure OpenAI.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

EWA Analyzer CLI

Turn a supported SAP EarlyWatch Alert report into a structured Excel review workbook.

An experimental, AI-assisted first-pass review tool for SAP Basis engineers, architects, and customer-facing technical teams, and a practical example of an agentic document-analysis workflow built with Python, LangGraph, Azure OpenAI, Pydantic, and openpyxl.

Human review required. Findings and recommendations are model-generated, not verified SAP advice. The CLI runs locally, but sends report content to your configured Azure OpenAI resource. Use only reports you are authorised to process. Model calls can incur charges.

What you get

One supported .doc report in; one .xlsx workbook out:

Workbook view What it helps you review
Executive Summary Model-generated narrative, proposed priorities, and status cautions
Section findings Evidence, potential impact, severity, and suggested investigations/actions
Cross-References Proposed relationships between findings, when available
Remediation Plan Suggested actions consolidated across sections
Coverage Which eligible sections were analysed, failed, or skipped
Analysis Trace Invocation metadata and representative source checks
Document Structure The report's section hierarchy
Token Usage & Cost Usage by deployment/phase and estimated cost when rates are configured

The workbook uses colour-coded severities, wrapped text, filters, frozen headers, and navigation links. Narrative/display fields render Markdown and supported HTML as readable text. Raw trace excerpts and validation queries remain literal, and model text is not executed as Excel formulas.

Requirements and supported input

  • Python 3.12 or newer and Git.
  • An Azure OpenAI resource, API key, and model deployment(s) compatible with the Responses API and structured outputs used by the Azure LangChain adapter.
  • Deployment access/quota for both configured roles: orchestration and specialist analysis. The same compatible deployment can be assigned to both roles; use its actual Azure deployment name.
  • A local filesystem that supports hard links for safe no-overwrite publication. Local NTFS works; unsupported/network filesystems may refuse publication.

Supported: Word 2003 XML content saved with a .doc suffix, as used by the supported EWA export format.

Not supported: binary Word .doc, .docx, PDF, scanned images, or arbitrary HTML uploads. Changing a binary Word/PDF filename to .doc does not convert its contents. If your export format differs, obtain a compatible export or perform conversion separately before using this CLI. Microsoft Word is not required to run the supported XML conversion.

Quick start

1. Clone and install

Windows PowerShell

git clone https://github.com/senjoyee/ewa-ai-analyzer-cli.git
cd ewa-ai-analyzer-cli
py -3 -m venv .venv
.\.venv\Scripts\python.exe -m pip install --upgrade pip
.\.venv\Scripts\python.exe -m pip install .
.\.venv\Scripts\ewa-analyzer.exe --help

Check that py -3 --version reports Python 3.12+. The explicit executable paths avoid PowerShell activation-policy issues and accidentally running a different installation.

macOS / Linux shell

git clone https://github.com/senjoyee/ewa-ai-analyzer-cli.git
cd ewa-ai-analyzer-cli
python3 -m venv .venv
.venv/bin/python -m pip install --upgrade pip
.venv/bin/python -m pip install .
.venv/bin/ewa-analyzer --help

Check that python3 --version reports Python 3.12+. Development verification has been performed on Windows; test your own OS and filesystem before relying on the workflow.

2. Configure Azure OpenAI

For a fresh clone, copy the placeholder environment file. Do not overwrite an existing .env when updating.

# Windows
Copy-Item .env.example .env
# macOS / Linux
cp .env.example .env

Edit .env with your resource settings:

AZURE_OPENAI_ENDPOINT=https://YOUR-RESOURCE.openai.azure.com/
AZURE_OPENAI_API_KEY=replace-with-your-key
AZURE_OPENAI_API_VERSION=2025-03-01-preview
ORCHESTRATOR_MODEL=your-orchestrator-deployment
SPECIALIST_MODEL=your-specialist-deployment
Variable Meaning
AZURE_OPENAI_ENDPOINT Azure resource root URL, not a URL ending in /openai/v1 or /responses
AZURE_OPENAI_API_KEY Your Azure resource key; never commit or share it
AZURE_OPENAI_API_VERSION API version compatible with your resource and deployments; the example shows the CLI default
ORCHESTRATOR_MODEL Actual Azure deployment name used for planning, cross-referencing, and synthesis
SPECIALIST_MODEL Actual Azure deployment name used for section analysis

Deployment aliases are not interchangeable with model names: a model being available in Azure does not mean it is deployed in your resource or supports the required capabilities.

The CLI reads .env from the current working directory, not from beside the installed package. Exported environment variables take precedence over .env values. If you run elsewhere, either put an approved .env there or export the variables in that shell. .env is ignored by Git.

3. Try the synthetic report

This command calls Azure OpenAI and may incur charges. The input is fabricated test data, not a customer report or accuracy benchmark.

# Windows
.\.venv\Scripts\ewa-analyzer.exe analyze --doc tests/fixtures/synthetic-ewa.doc --output output/synthetic-review.xlsx
# macOS / Linux
.venv/bin/ewa-analyzer analyze --doc tests/fixtures/synthetic-ewa.doc --output output/synthetic-review.xlsx

Open the resulting workbook and inspect Executive Summary, Coverage, Analysis Trace, and the section findings. To check the software without Azure credentials or model charges, run the offline test suite instead.

4. Analyse your own report

Use a supported report that you have permission to send to your Azure resource:

.\.venv\Scripts\ewa-analyzer.exe analyze --doc "C:\Reports\report.doc" --output "C:\Reviews\review.xlsx" --verbose

On macOS/Linux, use .venv/bin/ewa-analyzer and your own paths. Relative paths are resolved from the current directory. The CLI creates the output directory when needed; intermediate conversion files are kept in temporary storage, not beside your source report.

Existing output is refused unless you explicitly add --overwrite. Prefer a new filename when comparing runs. Input/output aliases and symlinked output paths are also refused.

Command reference

ewa-analyzer analyze --doc INPUT.doc --output REVIEW.xlsx [--config SETTINGS.yaml] [--overwrite] [--verbose]
Option Purpose
--doc Required, supported Word 2003 XML .doc input
--output Required, destination workbook
--config Explicit YAML settings/pricing file; not automatically discovered
--overwrite Allow replacing an existing output
--verbose Additional progress information; not raw provider-error contents

The equivalent module command is python -m ewa_pipeline analyze ..., using the virtual environment's Python.

Optional settings and pricing

Copy config.example.yaml to config.yaml and pass it explicitly:

.\.venv\Scripts\ewa-analyzer.exe analyze --doc tests/fixtures/synthetic-ewa.doc --output output/review-with-settings.xlsx --config config.yaml

The example includes request timeout, retry count, maximum parallelism, and optional pricing. Prices are USD per one million tokens, keyed by the actual Azure deployment name. Supply only verified rates. Without rates, cost is shown as unavailable, not zero; even configured cost figures are estimates, not an Azure bill.

Keep YAML non-secret and use environment variables for credentials. If you deliberately supply an azure_openai section in YAML, its populated values take precedence over environment values; missing values are filled from the environment. With no --config, no YAML file is loaded.

Exit codes and partial results

Exit code Meaning
0 Workbook published without a reported degraded-run condition; this does not establish correctness
1 Fatal input/configuration/conversion/model/publication error; no new workbook published
2 Partial/degraded workbook published, or a Click command-usage error; read the console message

For a partial workbook, inspect the coverage and trace banners before using the findings. Partial section coverage, correlation failure, or synthesis failure makes overall health Unknown. Partial trace evidence is reported separately and does not by itself change model-assessed health.

Coverage completeness means eligible sections were accounted for, not that every issue was found. A matched excerpt only confirms that the excerpt occurs in the section source. It does not validate the interpretation or prove a complete review. Validation flags are advisory heuristics, not guarantees about unflagged recommendations.

The agentic workflow

flowchart TD
    D["Supported EWA .doc"] --> N["Normalise to semantic HTML"]
    N --> T["Build section tree"]
    T --> P["Planning agent"]
    C["Skill catalogue"] -.-> P
    P --> V["Validate plan and backfill missing sections"]
    V --> A1["Section analyst 1"]
    V --> A2["Section analyst 2"]
    V --> AN["Section analyst N"]
    K["Skill instructions + selected references"] -.-> A1
    K -.-> A2
    K -.-> AN
    A1 --> M["Merge typed findings and traces"]
    A2 --> M
    AN --> M
    M --> X["Cross-reference agent"]
    X --> R["Validate referenced finding IDs"]
    R --> S["Synthesis"]
    M -.-> S
    S --> E["Deterministic Excel generation"]
    E --> H["Human review"]
Loading
  • Deterministic preparation: XML conversion and section indexing happen before model analysis.
  • Validated planning: invalid/duplicate tasks are removed and omitted eligible sections receive fallback tasks.
  • Focused parallel agents: LangGraph dispatches section analysts with task-specific context. Pydantic structures the returned findings.
  • Reusable skills: the packaged EWA skill and modular references cover performance, memory, database, security, batch, operations, remediation, and correlation guidance. Reference selection has a fallback and context-size limit.
  • Structured collaboration: findings receive stable section-prefixed IDs; downstream correlation and synthesis consume the structured outputs.
  • Deterministic delivery: openpyxl generates the workbook; models do not write or execute spreadsheet code.

This is an orchestrated analysis workflow, not an autonomous system administrator. It does not connect to SAP systems or execute suggested remediation. Packaged guidance must still be checked against current SAP documentation and the customer's context.

Troubleshooting

Symptom What to check
Unsupported .doc format The contents must be Word 2003 XML; a binary Word file or renamed .docx will not work
Configuration invalid/incomplete Check the required variables, current-directory .env, deployment names, and explicit YAML path; do not post credentials
Azure request or model failure Confirm the resource-root endpoint, API version, deployment capabilities, quota, permissions, and network access
Output already exists Choose another filename, or intentionally use --overwrite
Publication refused on a network drive Try a local filesystem with hard-link support; refusal avoids unsafe partial-copy publication
Exit code 2 with a workbook Review Coverage and Analysis Trace; do not treat the result as a verified healthy system
Updated source but old behaviour Reinstall the package into the environment whose CLI you are actually running
Cost says unavailable Add verified per-deployment rates in YAML and pass --config; inference is not necessarily free

Updating an installation

After pulling source updates, reinstall with the same virtual environment:

git pull
.\.venv\Scripts\python.exe -m pip install .
.\.venv\Scripts\python.exe -c "import ewa_pipeline.report.excel_generator as g; print(g.__file__)"

On macOS/Linux, substitute .venv/bin/python. Normal installation copies the package into site-packages; editing src alone does not update that installed copy. Existing workbooks are never rewritten by installation. Rerun analysis from the original supported document to produce new output.

Development and offline tests

The tests use synthetic data and fake model responses, with an offline transport regression for model request construction. They require no real Azure credentials and are not analytical-accuracy or performance benchmarks.

From the checkout, after installing dependencies:

# Windows
.\.venv\Scripts\python.exe -m unittest discover -s tests -v
# macOS / Linux
.venv/bin/python -m unittest discover -s tests -v

The test helper selects source-checkout code. Run suites sequentially: an existing workbook-test helper uses a shared temporary filename. A symlink-related test may be skipped when the OS does not permit symlink creation.

For editable development, use python -m pip install -e . with the environment's Python. For packaging checks, build a wheel, install it into a separate environment, and verify it outside the checkout so source imports cannot mask missing packaged skills.

Contributing and responsible use

Issues and pull requests are welcome. Include a minimal synthetic reproduction, Python/OS versions, the relevant command with sensitive paths removed, and the observed behaviour. Never attach customer EWA reports, generated customer workbooks, credentials, or unredacted provider payloads to public issues.

This is experimental software supplied as-is. It is not an official SAP product and makes no SAP accuracy, production-readiness, or guaranteed time-saving claim. Validate every recommendation before changing a system.

License and attribution

MIT License, copyright Joyee Sen. This standalone CLI was extracted from EWA AI Analyzer, based on commit 3909796, and retains its MIT license. This repository focuses on local document-to-workbook analysis rather than the original web application.

About

Agentic first-pass review of SAP EarlyWatch Alert reports: a Python/LangGraph CLI that turns supported Word XML reports into Excel workbooks using Azure OpenAI.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages