AgentShield · VulnSignal · LLM Security Eval Lab · DetectionForge · RiskOS · Frontier Agent Evals
Python SQL XGBoost scikit-learn Streamlit FastAPI NetworkX Docker
| Detect Abuse patterns Security anomalies Attack behavior Emerging campaigns |
Understand Graph relationships Behavioral context Threat provenance Attack paths |
Decide Risk calibration Policy thresholds Intervention ranking Operational tradeoffs |
Measure False positives Experiment impact Model drift Safety guardrails |
The projects below focus on the full analytical lifecycle: telemetry → features → models → explanations → decisions → measurement. Many combine supervised ML, anomaly detection, graph analytics, NLP, causal analysis, and analyst-facing dashboards rather than treating prediction as the final output.
🛡️ AgentShieldAI Agent Runtime Security Deterministic tool authorization, learned trajectory-risk evidence, human approval, redaction, and measurable risk-versus-friction controls.
|
AI Security Finding Intelligence End-to-end finding quality, model routing, deduplication, triage experiments, remediation outcomes, and verified resolution.
|
|
Adversarial AI Evaluation Prompt-injection, secret-leakage, hallucination, tool-authorization, approval-control, and response-trace evaluation.
|
Detection Engineering as Code Versioned detections, KQL compilation, malicious/benign replay, deterministic release gates, ML alert ranking, and active learning.
|
◇ RiskOSCalibrated Security Decisioning Risk decisions separated from prediction so thresholds and interventions can reflect exposure, capacity, false-positive cost, and policy constraints.
|
Long-Horizon Agent Evaluation Outcome quality, process safety, intervention behavior, efficiency, and calibration for agents that reason and act over multiple steps.
|
| Project | Focus |
|---|---|
| GoldenSet Factory | Production feedback → governed, deduplicated, versioned evaluation sets |
| EvalDrift | Offline-to-production divergence, confidence intervals, sustained alerts, investigation ranking |
| ReviewerIQ | Capacity-constrained human-review optimization with clean policy comparisons |
| FeatureDecay Lab | Signal decay, uncertainty, change points, replacement discovery, time-aware evaluation |
| DecisionStream | Leakage-safe event-time features, policy actions, replay, and champion/challenger evidence |
Shared engineering baseline: synthetic-data boundaries · Python 3.10–3.12 CI · reproducible manifests · benchmark artifacts · MIT licensing · security reporting.
These projects explore a recurring product question: how do you stop harmful behavior while minimizing friction for legitimate users? The emphasis is therefore not only recall, but precision, calibration, false-positive cost, exposure, intervention design, and ecosystem-level measurement.
| Project | Focus |
|---|---|
| MessageShield | RCS/RBM spam, phishing, impersonation, graph abuse signals, experiments, counterfactual analysis |
| CallShield AI | Scam calls, robocalls, voice impersonation, behavioral ML, anomaly detection, explainability |
| DeepTrace | Content authenticity, provenance, semantic similarity, coordinated narrative clustering |
| RiskOS | Calibrated Trust & Safety decisioning, exposure modeling, policy optimization |
| SignalForge | Fraud, abuse, and operational-risk ML with graph signals and cost-aware decisions |
| PhishGraph AI | Phishing/BEC detection using identity, URL, behavioral, and campaign-graph signals |
Methods represented: classification · gradient boosting · anomaly detection · graph features · clustering · calibration · SHAP-style explainability · A/B testing · counterfactual analysis.
This portfolio area treats AI systems as active security principals. The core questions include what an agent can access, what tools it can invoke, whether its trajectory remains policy-compliant, how safeguards fail under adversarial pressure, and how model behavior can be evaluated beyond single-turn accuracy.
| Project | Focus |
|---|---|
| AgentShield | Runtime AI-agent security, tool policy, trajectory risk, safeguard measurement |
| AegisMesh | Multi-agent security orchestration with policy-gated tools |
| AgentAtlas | AI-agent inventory, effective access, permission drift, anomaly prioritization |
| LLM Security Evaluation Lab | Prompt injection, leakage, misuse, tool authorization, intervention evaluation |
| Frontier Agent Evals | Long-horizon agent outcomes, process safety, efficiency, and calibration |
| Model Containment Eval Lab | Shutdown compliance, containment, tripwires, trace-risk monitoring |
| Oversight Integrity Lab | Counterfactual testing for compromised oversight and monitor effectiveness |
| VulnSignal | AI vulnerability-finding quality, model routing, validation, remediation outcomes |
Themes: agent authorization · prompt-injection resilience · containment · long-horizon evaluation · oversight integrity · model routing · tool-use policy · trajectory monitoring.
These systems focus on the operational path from high-volume security telemetry to an analyst decision: normalization, feature generation, detection logic, replay/testing, prioritization, investigation context, ATT&CK mapping, and measurable false-positive reduction.
| Project | Focus |
|---|---|
| DetectionForge | Detection-as-code, replay metrics, release gates, ML alert prioritization |
| Agentic SOC Investigator | Evidence-grounded SOC investigation, enrichment, ATT&CK mapping |
| InsiderGuard | UEBA, DLP, peer baselines, insider risk, exfiltration-sequence detection |
| Security Telemetry Lakehouse | Event normalization, dedupe, watermarking, feature windows, explainable detections |
| BrowserGuard | Browser-extension risk, session exposure, permission drift, hybrid anomaly detection |
| MacSentinel | macOS threat analytics, provenance graphs, streaming ML, robustness testing |
| Cybersecurity Analytics & AI | Identity, endpoint, network, SIEM, AI-safety, OSINT, and purple-team analytics |
Operational patterns: detection-as-code · UEBA · peer baselines · streaming features · SIEM analytics · evidence enrichment · ATT&CK mapping · analyst prioritization.
| Project | Focus |
|---|---|
| CodeSentinel AI | AI-native secure code review, CWE mapping, finding validation, remediation tracking |
| AdversarialWeb | Bots, scraping, credential stuffing, account takeover, detection-gap analysis |
This category connects application-layer evidence with ML-assisted prioritization: code findings, attack behavior, account-abuse patterns, validation, and remediation outcomes.
Rather than score cloud findings independently, these projects model relationships and reachability: which identities connect to which resources, how permissions create paths, what compromise can reach next, and which defensive intervention most efficiently reduces blast radius.
| Project | Focus |
|---|---|
| AttackPath AI | Cross-source identity and agentic attack-path detection |
| SaaSGraph | OAuth/SaaS exposure, token behavior, anomaly detection, blast radius |
| Counterfactual Security Engine | Attack-path simulation and defensive intervention ranking |
| CloudRescue | Cloud ransomware recovery assurance, recoverability gates, restore forecasting |
| InfraGuard AI | Critical-infrastructure AI assurance, provenance, safety envelopes, degraded-safe operation |
| GPU Trust Guardian | GPU trust, attestation, workload behavior, attack paths, analyst-agent guardrails |
| AI Data Center Security Digital Twin | Cyber-physical attack paths, blast radius, multivariate infrastructure telemetry ML |
Methods represented: graph traversal · attack-path ranking · anomaly detection · counterfactual intervention · recovery forecasting · attestation · digital twins · multivariate telemetry analysis.
These projects emphasize evidence quality and provenance. The goal is to preserve where an assertion came from, connect indicators and behaviors into usable context, expose contradictions, and make machine-generated recommendations inspectable by an analyst.
| Project | Focus |
|---|---|
| Threat Intelligence Knowledge Graph | CTI evidence paths, graph retrieval, analyst-gated link prediction |
| OSINT Threat Intelligence Agent | IOC investigation, provenance, contradictions, ATT&CK context |
| SupplyChain Guardian AI | SBOM, CI/CD, build provenance, dependency risk, policy release gates |
Themes: knowledge graphs · IOC enrichment · provenance · ATT&CK context · SBOM analysis · build integrity · dependency risk · analyst-in-the-loop reasoning.
| Layer | Techniques |
|---|---|
| Supervised ML | Logistic regression, gradient boosting, XGBoost, multiclass classification, ensemble scoring |
| Unsupervised ML | Isolation Forest, clustering, behavioral baselines, novelty detection |
| NLP | Text classification, TF-IDF, semantic similarity, transcript analysis, threat-text enrichment |
| Graph analytics | Attack paths, knowledge graphs, campaign graphs, blast radius, centrality, reachability |
| Model evaluation | PR-AUC, ROC-AUC, precision/recall, calibration, false-positive guardrails, challenger/champion comparison |
| Explainability | Feature importance, local risk explanations, evidence paths, analyst-readable reasoning |
| Experimentation | A/B testing, guardrail metrics, counterfactual analysis, intervention measurement |
| Decision science | Threshold optimization, exposure modeling, cost-aware decisions, capacity-aware policy |
Raw telemetry
│
▼
Normalization & quality
│
▼
Behavioral / graph / text features
│
▼
Rules + ML + anomaly detection
│
▼
Calibrated risk
│
├────────► Explanation / evidence
│
▼
Policy or analyst decision
│
▼
Outcome measurement
│
▼
Feedback, tuning & monitoring
The common design principle across the portfolio is that a model score is not the end product. Useful security systems also need evidence, thresholds, guardrails, operational context, and a way to measure whether the intervention actually improved the outcome.
| Data & ML Python SQL pandas scikit-learn XGBoost PySpark |
Security analytics Splunk / SIEM MITRE ATT&CK Identity telemetry Endpoint telemetry Threat intelligence Detection engineering |
AI systems LLMs RAG Agent workflows Evaluation harnesses Tool policy Knowledge graphs |
Productization Streamlit FastAPI Docker Dashboards Model monitoring Experimentation |
Earlier data-science and coursework repositories remain available for historical context. The six projects above represent the current cybersecurity and AI-security focus.
Data Science Projects · Python Proficiency Test · Social Media Mining · NLP Projects · Customer Marketing Analytics · Electronic Arts Data Science Test
Security-first — optimize for adversarial behavior, uncertainty, false-positive cost, and evidence quality.
Explainable by design — pair risk scores with features, graph paths, provenance, or analyst-readable evidence.
Product-aware — connect model outputs to interventions, user friction, capacity constraints, and measurable outcomes.
Evaluation-driven — use holdouts, replay, guardrails, calibration, experiments, and counterfactuals instead of relying on headline accuracy.
Systems-oriented — build the surrounding data, decision, and monitoring layers rather than presenting isolated notebooks.
