Skip to content

Latest commit

 

History

23 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Lineage-Aware Memory Governance for Enterprise AI Agents

Paper: Lineage-Aware Memory Governance: A Derivation-Gated Framework for Privacy-Preserving Column-Level Access Control in Enterprise AI Agents Authors: Venkata Sangaraju · Sudhir Vissa (SAGE7 AI) Published: IEEE Access, Early Access, open access · DOI: 10.1109/ACCESS.2026.3730363 · also on IEEE Xplore Reference implementation: amu-governance (MIT, pip-installable: pip install amu-governance) · Zenodo software DOI


Overview

Shared memory across AI agents offers efficiency gains in enterprise analytics but introduces two underexplored failure modes:

  1. Column-level data leakage — an agent retrieves an insight derived from columns it is not permitted to see.
  2. Silent metric-definition conflicts — two teams compute the same KPI through divergent derivation paths and the wrong definition propagates silently.

This repository contains the full simulation code, experiment scripts, result data, and paper source for a proposed solution: the Analytical Memory Unit (AMU) schema and lineage-gated retrieval policy.


Repository Structure

Lineage-Aware-Memory/
│
├── model.py                    # AMU dataclass, Lineage, LineageStep, schema definitions
├── systems.py                  # NoMemorySystem, NaiveMemorySystem, LineageAwareSystem
├── simulate.py                 # Event generator, run_simulation(), run_seed_sweep(30)
│
├── gen_figures.py              # Generates all paper figures (matplotlib)
├── tpch_experiment.py          # TPC-H 8-table schema, 30-seed sweep
├── fuzzy_experiment.py         # Fuzzy conflict-detection study (D1/D2/D3 + extended)
├── degradation_experiment.py   # Lineage-completeness degradation sweep
├── benchmark_runtime.py        # Empirical runtime benchmarks (µs-level)
├── stats_analysis.py           # Bootstrap CI + Mann-Whitney U significance tests
│
├── results.csv                 # Single seed=42 simulation results
├── sweep_summary.csv           # 30-seed sweep summary (means ± std dev)
├── sweep_raw.json              # 30-seed sweep, full per-seed raw values (feeds stats_analysis.py)
├── tpch_results.json           # TPC-H experiment: per-seed raw data + summary
├── fuzzy_results.json          # Original 10-pair fuzzy detection results
├── fuzzy_results_extended.json # Extended 43-pair fuzzy study with bootstrap CI
├── degradation_results.json    # Completeness sweep: leak rate vs. reporting fraction
├── benchmark_results.json      # Runtime measurements in µs across operation types
├── stats_results.json          # Bootstrap CI + Mann-Whitney p-values
├── adversarial_lineage_experiment.py  # Exp. 7: materialization-boundary leak + closure fix
├── adversarial_lineage_results.json   # Exp. 7 output
├── storage_overhead_measurement.py    # Empirical check of the paper's 4-8x storage estimate
├── storage_overhead_results.json      # Measurement output
├── test_package_conformance.py        # Root modules vs. amu-governance package, 30-seed check
├── RESULTS.md                  # Honest assessment of what the results do/don't show
│
├── agent_demo/
│   ├── demo.py                 # End-to-end walkthrough: SQLite DB + two agent functions
│   └── transcript.md           # Auto-generated readable case-study transcript
│
├── figures/                     # Output of gen_figures.py — reproducibility
│   ├── fig1_comparison_bars.png/.pdf   # Per-metric bar chart
│   ├── fig2_architecture.png/.pdf      # AMU architecture diagram
│   ├── fig3_tradeoff.png/.pdf          # Governance/efficiency scatter
│   ├── fig4_tpch_comparison.png/.pdf   # TPC-H vs synthetic comparison
│   └── fig5_degradation.png/.pdf       # Leak rate vs. lineage completeness
│
└── extension_paper/             # Follow-on paper: new experiments, not in the
                                  # published IEEE Access paper — see below and
                                  # extension_paper/README.md for full detail
    ├── src/amu_ext/              # Library code: D4/D5 detectors, aggregation
    │                             #  schema, materialization closure, SQL
    │                             #  extraction, stats helpers
    ├── data/                     # conflict_dataset_43.json (ported),
    │                             #  conflict_dataset_ao_10.json (new AO category)
    ├── experiments/              # run_*.py — full-statistical-power scripts
    ├── tests/                    # unit / integration / regression tiers
    ├── docs/                     # threat_model_extension.md, limitations.md
    └── results/                  # RESULTS.md + timestamped run folders

Relationship between this repository and the amu-governance package

The mechanism appears in four places in this project, and they are independent implementations that happen to agree, not one shared codebase:

  1. model.py / systems.py in this repo — used by simulate.py, degradation_experiment.py, and benchmark_runtime.py. Hardcodes the paper's 5-table schema and department permissions as module globals.
  2. Self-contained reimplementations embedded directly in tpch_experiment.py (TAMU/TLineage/TLineageAwareSystem, TPC-H schema) and fuzzy_experiment.py — neither imports model.py/systems.py.
  3. amu-governance on PyPI — a policy-agnostic rewrite (GovernancePolicy replaces the hardcoded globals) intended as the general-purpose reference implementation for reuse outside this paper's specific schema.

This split is a known, low-risk-but-real maintenance duplication, not a sign of drift: an equivalence check (test_package_conformance.py, seeded, 30 runs) confirms model.py/systems.py and amu-governance produce identical served/leaked/reused/blocked/conflict-flagged outcomes on the exact workload behind the paper's headline synthetic-schema numbers. The root modules here are the frozen artifact that produced the published results; amu-governance is the maintained, schema-agnostic descendant for new work — including adversarial_lineage_experiment.py above, which is written against the package rather than the root modules for that reason. Consolidating the TPC-H/fuzzy reimplementations onto one shared codebase would reduce this duplication further but was intentionally left out of scope here, since it would require re-running and re-verifying every published number rather than adding new, additive experiments.


Reproducing the Results

Install dependencies first:

pip install -r requirements.txt

All experiments are deterministic (seeded). Run in order:

# 1. Core simulation — reproduces results.csv and sweep_summary.csv
python simulate.py

# 2. TPC-H schema experiment — reproduces tpch_results.json
python tpch_experiment.py

# 3. Fuzzy conflict-detection study — extended 43-pair dataset
#    Produces fuzzy_results_extended.json (fuzzy_results.json unchanged)
python fuzzy_experiment.py

# 4. Lineage-completeness degradation sweep
#    Tests what happens when agents under-report their derivation columns
#    Completeness levels: 0.5, 0.75, 0.9, 0.95, 1.0 × 30 seeds × 2 schemas
python degradation_experiment.py

# 5. Runtime benchmark (2000 iterations per configuration)
#    Measures sensitivity-tag derivation, definition-hash, retrieve (n=1–50),
#    write/conflict-detection — all in µs
python benchmark_runtime.py

# 6. Statistical significance testing
#    Bootstrap 95% CI + Mann-Whitney U on all naive-vs-lineage-aware comparisons
python stats_analysis.py

# 7. Regenerate all paper figures (including new Figure 5: degradation)
python gen_figures.py

# 8. Agent demo — end-to-end walkthrough with a real SQLite database
#    Finance + Marketing agents through LineageAwareSystem
#    Generates agent_demo/transcript.md
python agent_demo/demo.py

# 9. Adversarial lineage experiment (post-publication addendum, not in the
#    IEEE Access paper) — materialization-boundary leak + closure fix.
#    Requires the amu-governance package: pip install amu-governance
python adversarial_lineage_experiment.py

Key Results

Core Simulation (Paper §5.1–5.2)

Experiment Naive Shared Memory Lineage-Aware Significance
Leak rate — synthetic schema (30 seeds) 18.8% ± 3.7% 0.0% ± 0.0% p < 0.001
Leak rate — TPC-H schema (30 seeds) 25.5% ± 4.5% 0.0% ± 0.0% p < 0.001
Memory reuse — synthetic 95.8% ± 0.0% 82.6% ± 2.3% p < 0.001
Memory reuse — TPC-H 95.8% ± 0.0% 81.5% ± 4.8% p < 0.001
Conflict recall 0.0% 100.0% —

All comparisons: Mann-Whitney U, p < 0.001 (bootstrap 95% CI excludes zero). The 0% leakage result is a formal guarantee (Theorem 1), not an empirical finding.

Post-publication addendum: adversarial lineage testing (not in the paper)

adversarial_lineage_experiment.py — the paper's own stated #1 roadmap priority, added after publication. Tests the gate against a materialization boundary: a sensitive metric (risk_score, derived from income) is persisted as a column, then a second metric is derived from that column alone, with lineage that never names income. Full writeup in RESULTS.md.

Gate Leak rate (30 seeds) Block rate
Stock LineageAwareSystem (published, unmodified) 32.7% ± 9.3% 0.0%
+ transitive lineage closure (candidate fix) 0.0% ± 0.0% 32.7%

Fuzzy Conflict Detection — Extended Study (43 pairs: TC/TV/LO/CS)

Detector Precision Recall F₁ 95% CI (F₁) LO category
D1 — Exact hash 0.651 1.000 0.789 [0.677, 0.883] TP: 8/8
D2 — Jaccard (τ=0.70) 0.676 0.821 0.742 [0.604, 0.853] TP: 3/8
D3 — Column-graph 1.000 0.714 0.833 [0.703, 0.933] FN: 8/8

New finding (LO category): D3's perfect precision/recall on the original 10 pairs does not generalise. Logic-operator changes (AND→OR) with identical column sets are true semantic conflicts that D3 misses entirely. D1 catches all 8 LO cases; D2 catches 3/8. A production deployment needs D3 + filter-logic check.

Lineage-Completeness Degradation (New — degradation_experiment.py)

Schema Completeness Leak Rate (mean ± sd) vs. full lineage
Synthetic 0.50 16.3% ± 3.3% p < 0.001*
Synthetic 0.75 0.0% ± 0.0% not sig.
Synthetic 1.00 0.0% ± 0.0% baseline
TPC-H 0.50 10.3% ± 4.7% p < 0.001*
TPC-H 0.75 0.9% ± 1.1% p < 0.001*
TPC-H 1.00 0.0% ± 0.0% baseline

At completeness=0.50, incomplete lineage reporting causes ~10–16% leak rate — matching the naive baseline. TPC-H is more sensitive (sensitive columns spread across more join steps). The safety guarantee requires completeness=1.0 (Assumption 1 in the paper).

Runtime Benchmarks (New — benchmark_runtime.py, 2000 iterations)

Operation Mean (µs) Notes
Sensitivity-tag derivation 0.36 O(k·c), cached at write time
Definition-hash (SHA-256) 1.20 O(k·c), cached at write time
Gate-check predicate (set diff) 0.09 Sub-microsecond
Retrieve — best-case (n=1–50) ~0.64 O(1): first candidate safe
Retrieve — worst-case n=1 0.66 All blocked: linear scan begins
Retrieve — worst-case n=50 13.82 21× vs n=1 — confirms O(n)
Write/conflict-detect (early exit) 2.37 O(1): conflict found on first pair
Write/conflict-detect (full scan, n=50) 1.63 O(n): no conflict, scans all

O(n) worst-case scaling confirmed empirically. At n ≤ 5 (realistic production), retrieve worst-case is < 2 µs — negligible against any real analytics query (typically ms–seconds).


Follow-on Work: Extension Paper (PeerJ Computer Science / arXiv)

extension_paper/ contains a second, in-progress paper built on top of the published IEEE Access work — new experiments only, nothing here changes the published paper's LaTeX/PDF. See extension_paper/README.md and extension_paper/results/RESULTS.md for the full detail; in short, three groups of new work:

  • Group A — detector evaluation. The original paper proposed a fourth conflict detector, D4 (structural gate + filter-logic operator-topology diff), but never evaluated it. This adds D4 and D5 (D4 + an aggregation-equality check), evaluated against the original 43-pair dataset plus a new Aggregation-Operator (AO) category, with bootstrap 95% CIs and a paired McNemar test.
  • Group B — materialization-boundary threat. Formalizes and generalizes the transitive-lineage-closure threat (see the adversarial addendum above) across multiple chain/topology configurations, with a regression test proving the closure fix does not silently over-claim on an unregistered materialization edge.
  • Group C — reproducibility and correctness hardening. Wires the package-conformance check into CI, re-measures storage overhead, adds a before/after demonstration of the SQL unqualified-column bugfix, and fixes a diagnostic-flag bug in Algorithm 1's retrieval gate (root repo and amu-governance package).

Reproduce with cd extension_paper && make setup && make all — see that directory's README for the full command reference and test-tier breakdown.


Paper

This repository holds the code, data, and reproducible experiments behind the paper — not the manuscript itself. Read the published version via the links at the top of this README (open access on IEEE Access / IEEE Xplore). figures/ contains the exact result figures the code here produces (regenerate with python gen_figures.py); Lineage-Aware-Memory-paper/ was previously tracked here but has been removed in favor of keeping this repo focused on runnable code and reproducible results.


Citation

@article{sangaraju2026lineage,
  author  = {Venkata Sangaraju and Sudhir Vissa},
  title   = {Lineage-Aware Memory Governance: A Derivation-Gated Framework
             for Privacy-Preserving Column-Level Access Control in
             Enterprise {AI} Agents},
  journal = {IEEE Access},
  year    = {2026},
  doi     = {10.1109/ACCESS.2026.3730363}
}

License

MIT — see LICENSE. This repository contains code, data, and experiment results only; the paper manuscript itself is published separately (see the link at the top of this README) and is not part of this repository.

About

Lineage-Aware-Memory

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages