Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 8 additions & 4 deletions docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -102,10 +102,14 @@ supply only `action` and render as one line.

## Troubleshooting

See `troubleshooting/service.py`. `gather_evidence()` always runs before `diagnose()`; the
platform never hands an LLM a one-line failure description and asks it to guess. V1 ships one
deliberate simulated failure (a Kubernetes readiness probe pointed at the wrong path) so the
mechanism is demonstrated end to end.
See `troubleshooting/service.py` and `workflows/troubleshooting_flow.py`. Evidence gathering
always precedes diagnosis: the platform never hands an LLM an unstructured failure description.
The troubleshooting engine manages bounded operational recovery scenarios (`port_conflict`,
`missing_config`, `health_check_failure`, `resource_limit`) with an explicit lifecycle:
`SETUP -> INJECT -> OBSERVE -> EXPLAIN -> REMEDIATE -> VERIFY -> CLEANUP`.
Observations (facts) are separated from interpretations (hypotheses), progressive hints (0-4)
guide the learner without preempting discovery, and recovery is deterministically re-verified
before completion.

## Persistence

Expand Down
13 changes: 11 additions & 2 deletions docs/roadmap.md
Original file line number Diff line number Diff line change
Expand Up @@ -67,14 +67,23 @@ and PR CI. See `docs/devsecops.md`.
The gate is now bound to each deployment candidate before any real apply. Live
Azure verification remains the final opt-in acceptance step.

## Milestone 4: realistic troubleshooting with recovery verification (done)

`devops-learn troubleshoot` introduces structured operational incident recovery
scenarios (`port_conflict`, `missing_config`, `health_check_failure`, `resource_limit`)
with an explicit lifecycle:
SETUP -> INJECT -> OBSERVE -> EXPLAIN -> REMEDIATE -> VERIFY -> CLEANUP.
Observations are strictly separated from interpretations; progressive assistance
(Levels 0-4) guides the learner without leaking answers; recovery is deterministically
re-verified; and results are honestly distinguished between `LIVE VERIFIED` and
`SIMULATED / TESTED`.

## Further out (not yet milestoned)

- Real AWS/GCP `CloudProvider` implementations.
- A bundled Go example project (detection already exists).
- A richer `ProjectAnalyzer` (dependency graph analysis, actual secret-scanning
integration, Helm/Kubernetes-manifest-aware analysis).
- More troubleshooting scenarios beyond the readiness-probe and container-exit
cases.
- A real Anthropic-backed explanation path tested end to end (V1 tests
`AnthropicProvider` only for import/credential behavior, not live calls).
- A web UI reusing the existing `workflows`/`Ui` boundary.
Expand Down
17 changes: 17 additions & 0 deletions docs/safety.md
Original file line number Diff line number Diff line change
Expand Up @@ -121,3 +121,20 @@ deliberately never returns a specific dollar figure, since it has no live
pricing data to draw from.
Cost impact in a `Recommendation` (e.g. "lower likely cost than a managed cluster") is always
qualitative for the same reason.

## Troubleshooting scenario safety and bounded fault injection

`devops-learn troubleshoot` creates realistic operational failure scenarios while strictly
preserving system integrity:

- **Bounded resource constraints**: Memory limit testing is confined to isolated container
configurations (e.g. 6MB container limits) or deterministic simulations; it never starves or
exhausts host machine memory.
- **Port isolation**: Port collisions use local ephemeral sockets/containers bound to loopback
(`127.0.0.1`) and are guaranteed to close during teardown.
- **Harmless mock configurations**: Missing configuration scenarios test application fail-fast
behavior using non-sensitive mock settings (e.g. `REQUIRED_CONFIG_KEY`), never real secrets.
- **Guaranteed teardown**: Every scenario runs cleanup inside a `finally` block to stop containers
and release occupied resources regardless of whether the exercise succeeds, fails, or is aborted.
- **Honest capability reporting**: Scenarios clearly declare whether execution was `LIVE VERIFIED`
or `SIMULATED / TESTED`. If Docker is absent, execution falls back cleanly to simulation.
103 changes: 103 additions & 0 deletions src/devops_learn/cli/commands/troubleshoot.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,103 @@
"""`devops-learn troubleshoot`: diagnostic reasoning and deterministic recovery verification."""

from __future__ import annotations

import argparse
from typing import Any

from devops_learn.bootstrap import Platform
from devops_learn.cli.terminal_ui import TerminalUi
from devops_learn.workflows.troubleshooting_flow import (
TroubleshootingOptions,
list_troubleshooting_scenarios,
run_troubleshooting_flow,
)


def register(subparsers: "argparse._SubParsersAction[argparse.ArgumentParser]") -> None:
parser = subparsers.add_parser(
"troubleshoot",
help="Troubleshooting scenarios with progressive assistance and recovery verification",
)
commands = parser.add_subparsers(dest="troubleshoot_command", required=True)

# list
list_cmd = commands.add_parser("list", help="List all available troubleshooting scenarios")
list_cmd.set_defaults(handler=run_list)

# run
run_cmd = commands.add_parser("run", help="Start and solve a troubleshooting scenario")
run_cmd.add_argument(
"scenario",
help="Scenario ID (e.g. port_conflict, missing_config, health_check_failure)",
)
run_cmd.add_argument(
"--hint-level",
type=int,
choices=[0, 1, 2, 3, 4],
default=None,
help="Progressive hint level (0=evidence, 1=inspection, 2=subsystem, 3=cause, 4=fix)",
)
run_cmd.add_argument(
"--remediation",
type=str,
default=None,
help="Proposed remediation action or key=value parameter",
)
run_cmd.add_argument(
"--real",
action="store_true",
help="Attempt real Docker/local tool execution if available",
)
run_cmd.add_argument(
"--simulate",
action="store_true",
help="Force simulated execution mode for offline/test environments",
)
run_cmd.add_argument(
"--path",
default=".",
help="Target project root directory",
)
run_cmd.set_defaults(handler=run_scenario)

# doctor
doc_cmd = commands.add_parser("doctor", help="Check environment readiness for troubleshooting")
doc_cmd.set_defaults(handler=run_doctor)


def run_list(args: argparse.Namespace, platform: Platform) -> None:
list_troubleshooting_scenarios(platform, TerminalUi())


def run_scenario(args: argparse.Namespace, platform: Platform) -> None:
remediation_params: dict[str, Any] = {}
if args.remediation and "=" in args.remediation:
for item in args.remediation.split():
if "=" in item:
k, v = item.split("=", 1)
remediation_params[k.strip()] = v.strip()

simulate: bool | None = None
if args.simulate:
simulate = True
elif args.real:
simulate = False

options = TroubleshootingOptions(
scenario_id=args.scenario,
hint_level=args.hint_level,
remediation_action=args.remediation,
remediation_params=remediation_params,
project_root=args.path,
simulate=simulate,
interactive=args.remediation is None and args.hint_level is None,
)
evidence = run_troubleshooting_flow(platform, TerminalUi(), options)
if not evidence.resolved and args.remediation is not None:
raise SystemExit(1)


def run_doctor(args: argparse.Namespace, platform: Platform) -> None:
from devops_learn.workflows.doctor_flow import run_doctor as execute_doctor
execute_doctor(platform, TerminalUi())
5 changes: 4 additions & 1 deletion src/devops_learn/cli/main.py
Original file line number Diff line number Diff line change
Expand Up @@ -23,6 +23,7 @@
terraform,
ai_test,
config,
troubleshoot,
)
from devops_learn.config.settings import load_settings
from devops_learn.domain.enums import ExecutionMode
Expand Down Expand Up @@ -61,6 +62,7 @@
report,
ai_test,
config,
troubleshoot,
)


Expand Down Expand Up @@ -215,7 +217,8 @@ def _tools_for_args(args: argparse.Namespace) -> dict[str, Tool] | None:
"security_policy": PolicyTool(),
"azure": AzureCliTool(),
}
if getattr(args, "real_tools", False) or args.command == "local":
is_troubleshoot_real = args.command == "troubleshoot" and getattr(args, "real", False)
if getattr(args, "real_tools", False) or args.command == "local" or is_troubleshoot_real:
return {
"python": RealPythonTool(),
"git": SimulatedGitTool(),
Expand Down
4 changes: 4 additions & 0 deletions src/devops_learn/domain/enums.py
Original file line number Diff line number Diff line change
Expand Up @@ -162,6 +162,10 @@ class AuditEventType(Enum):
DEPLOYMENT_FAILED = "deployment_failed"
TROUBLESHOOTING_STARTED = "troubleshooting_started"
DIAGNOSIS_PRODUCED = "diagnosis_produced"
TROUBLESHOOTING_REMEDIATION_ATTEMPTED = "troubleshooting_remediation_attempted"
TROUBLESHOOTING_VERIFIED = "troubleshooting_verified"
TROUBLESHOOTING_FAILED = "troubleshooting_failed"
TROUBLESHOOTING_COMPLETED = "troubleshooting_completed"
ROLLBACK_PERFORMED = "rollback_performed"
SESSION_COMPLETED = "session_completed"

Expand Down
95 changes: 95 additions & 0 deletions src/devops_learn/domain/troubleshooting_models.py
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,20 @@
from __future__ import annotations

from dataclasses import dataclass, field
from enum import IntEnum
from typing import Any, Mapping

from devops_learn.domain.learner_profile_models import CompetencyArea


class HintLevel(IntEnum):
"""Progressive assistance levels (0 to 4)."""

EVIDENCE = 0
INSPECTION = 1
SUBSYSTEM = 2
ROOT_CAUSE = 3
REMEDIATION = 4


@dataclass(frozen=True)
Expand All @@ -32,3 +46,84 @@ class Diagnosis:
explanation: str
recommended_fix: str
learning_moment: str | None = None


@dataclass(frozen=True)
class Observation:
"""Factual, deterministic tool output or system measurement."""

source: str # e.g. "docker.logs", "http_probe", "socket_status", "container_exit"
content: str
exit_code: int | None = None
is_error: bool = False
details: Mapping[str, Any] = field(default_factory=dict)


@dataclass(frozen=True)
class Interpretation:
"""Analytical deduction separated from raw observation."""

observation_summary: str
likely_subsystem: str
hypothesis: str
confidence: float = 1.0


@dataclass(frozen=True)
class RemediationAttempt:
"""Learner's proposed operational or configuration fix."""

scenario_id: str
action: str
parameters: Mapping[str, Any] = field(default_factory=dict)


@dataclass(frozen=True)
class VerificationResult:
"""Deterministic recovery verification outcome."""

success: bool
summary: str
observations: tuple[Observation, ...] = field(default_factory=tuple)
is_live: bool = False
details: Mapping[str, Any] = field(default_factory=dict)


@dataclass(frozen=True)
class TroubleshootingEvidence:
"""Complete ledger of troubleshooting investigation and recovery."""

scenario_id: str
before_state: tuple[Observation, ...]
remediation: RemediationAttempt | None = None
after_state: tuple[Observation, ...] = field(default_factory=tuple)
verification: VerificationResult | None = None
resolved: bool = False
mode_label: str = "(simulated)"


@dataclass(frozen=True)
class TroubleshootingScenario:
"""Specification of a bounded troubleshooting problem."""

scenario_id: str
title: str
learning_objective: str
category: CompetencyArea
fault_description: str
expected_symptoms: tuple[str, ...]
allowed_diagnostic_tools: tuple[str, ...]
hints: Mapping[int, str]
success_criteria: str
cleanup_requirements: str


@dataclass(frozen=True)
class TroubleshootingSession:
"""Active troubleshooting scenario lifecycle state."""

scenario: TroubleshootingScenario
is_live: bool
project_root: str
evidence: TroubleshootingEvidence
active: bool = True
1 change: 1 addition & 0 deletions src/devops_learn/troubleshooting/scenarios/__init__.py
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
"""Troubleshooting scenario package."""
51 changes: 51 additions & 0 deletions src/devops_learn/troubleshooting/scenarios/base.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,51 @@
"""Base interface for troubleshooting scenario handlers."""

from __future__ import annotations

from abc import ABC, abstractmethod
from dataclasses import dataclass, field
from typing import Any

from devops_learn.domain.troubleshooting_models import (
Observation,
RemediationAttempt,
TroubleshootingScenario,
VerificationResult,
)
from devops_learn.tools.service import ToolService


@dataclass
class ScenarioContext:
scenario: TroubleshootingScenario
is_live: bool
project_root: str
tool_service: ToolService
state: dict[str, Any] = field(default_factory=dict)


class ScenarioHandler(ABC):
@property
@abstractmethod
def definition(self) -> TroubleshootingScenario:
"""The declarative specification of this troubleshooting scenario."""

@abstractmethod
def setup_and_inject(self, context: ScenarioContext) -> tuple[Observation, ...]:
"""Perform setup, safely inject the fault, and return baseline observations."""

@abstractmethod
def remediate(
self, context: ScenarioContext, attempt: RemediationAttempt
) -> tuple[Observation, ...]:
"""Apply the remediation attempt and return intermediate observations."""

@abstractmethod
def verify(
self, context: ScenarioContext, attempt: RemediationAttempt
) -> VerificationResult:
"""Deterministically verify if recovery was achieved."""

@abstractmethod
def cleanup(self, context: ScenarioContext) -> None:
"""Guaranteed cleanup of temporary resources, sockets, or containers."""
Loading
Loading