# Implementation Plan ## Overview Bugfix implementation for five pipeline health issues: stuck parsed docs, missing price fallback, uncalibrated sentiment scores, idle signal-engine pods, and overly strict quality gate threshold. Tasks follow the exploratory bugfix workflow: explore bugs via tests, preserve existing behavior, implement fixes, validate. ## Tasks - [x] 1. Write bug condition exploration test - **Property 1: Bug Condition** - Pipeline Health Degradation - **CRITICAL**: This test MUST FAIL on unfixed code - failure confirms the bugs exist - **DO NOT attempt to fix the test or the code when it fails** - **NOTE**: This test encodes the expected behavior - it will validate the fix when it passes after implementation - **GOAL**: Surface counterexamples that demonstrate all five bugs exist - **Scoped PBT Approach**: Scope properties to the concrete failing cases for each bug condition - Test file: `tests/test_pbt_pipeline_health_bug_condition.py` - **Bug 1 - Batch Overflow**: Create 600 documents in `parsed` status older than 30 min, run `recover_stale_documents()`, assert up to 500 are recovered per cycle (will FAIL on unfixed code which caps at 100) - **Bug 2 - Price Fallback Gap**: Call `create_prediction_snapshot()` for ticker with no position and no exact market_snapshots match but data within 24h exists, assert `price_at_prediction` is NOT NULL (will FAIL on unfixed code which returns NULL) - **Bug 3 - Sentiment Bias**: Generate 50 impact records with systematic negative bias (mean=-0.3, stddev=0.1), compute weighted signals via aggregation, assert normalized output is zero-centered (will FAIL on unfixed code which passes raw scores) - **Bug 5 - Quality Gate Staleness**: Set most recent metric snapshot to 26h ago, evaluate quality gate, assert it passes (will FAIL on unfixed code which rejects at 24h) - Run tests on UNFIXED code - **EXPECTED OUTCOME**: Tests FAIL (this is correct - it proves the bugs exist) - Document counterexamples: batch capped at 100, price stored as NULL, sentiment heavily negative, quality gate returns `passed=False` - Mark task complete when tests are written, run, and failures are documented - _Requirements: 1.1, 1.2, 1.3, 1.4, 1.6_ - [x] 2. Write preservation property tests (BEFORE implementing fix) - **Property 2: Preservation** - Pipeline Behavior Unchanged for Non-Bug Inputs - **IMPORTANT**: Follow observation-first methodology - Test file: `tests/test_pbt_pipeline_health_preservation.py` - **Normal Doc Processing**: Observe that documents < 30 min old are never touched by `recover_stale_documents()` on unfixed code. Write property: for all documents with age < 30 min, recovery task does NOT enqueue them. - **Primary Price Path**: Observe that when `fetch_latest_close_price()` returns a valid price, no fallback is invoked. Write property: for all tickers where primary price exists, result equals primary price. - **Balanced Sentiment**: Observe that tickers with symmetric impact_score distributions (mean ≈ 0) produce neutral signals. Write property: for all impact_score sets with mean ≈ 0, normalized output remains centered around 0. - **Quality Gate Normal**: Observe that snapshots < 48h old meeting thresholds pass the quality gate. Write property: for all snapshot ages in [0, 48h) meeting metric criteria, quality gate returns `passed=True`. - **Failed Extraction Independence**: Observe `retry_failed_extractions()` behavior is unaffected. Write property: for all documents in `extraction_failed` status, retry logic unchanged. - Run tests on UNFIXED code - **EXPECTED OUTCOME**: Tests PASS (this confirms baseline behavior to preserve) - Mark task complete when tests are written, run, and passing on unfixed code - _Requirements: 3.1, 3.2, 3.3, 3.5, 3.6, 3.7_ - [x] 3. Fix: Signal Engine Scale Down (Helm values) - [x] 3.1 Set signal-engine replicas to 0 in all Helm values files - In `infra/helm/stonks-oracle/values.yaml`: change `signalEngine.replicas` from `1` to `0` - In `infra/helm/stonks-oracle/values-beta.yaml`: add/set `signalEngine.replicas: 0` under `services:` - In `infra/helm/stonks-oracle/values-paper.yaml`: add/set `signalEngine.replicas: 0` under `services:` - _Bug_Condition: input.signalEngineReplicas > 0 AND input.dualPipelineEnabled == FALSE_ - _Expected_Behavior: signalEngine.replicas == 0 when dual pipeline disabled_ - _Preservation: Signal-engine architecture remains functional when re-enabled with replicas > 0_ - _Requirements: 2.5, 3.4_ - [x] 4. Fix: Quality Gate Threshold Relaxation - [x] 4.1 Change max_snapshot_age_hours default from 24 to 48 - File: `services/trading/model_quality_gate.py` - In `QualityGateConfig` class, change `max_snapshot_age_hours: int = 24` to `max_snapshot_age_hours: int = 48` - _Bug_Condition: input.snapshotAgeHours > 24 AND input.snapshotAgeHours <= 48 AND input.maxSnapshotAgeConfig == 24_ - _Expected_Behavior: Quality gate accepts snapshots up to 48h old_ - _Preservation: Snapshots < 48h meeting criteria continue to promote to live_eligible_ - _Requirements: 2.6, 3.5_ - [x] 5. Fix: Stuck Parsed Docs Recovery - [x] 5.1 Lower STALE_PARSED_THRESHOLD_MINUTES from 240 to 30 - File: `services/scheduler/app.py` - Change constant: `STALE_PARSED_THRESHOLD_MINUTES = 30` - Documents stuck longer than 30 minutes are likely orphaned - _Requirements: 2.1_ - [x] 5.2 Increase recovery batch LIMIT from 100 to 500 - File: `services/scheduler/app.py` - In `recover_stale_documents()` SQL query, change `LIMIT 100` to `LIMIT 500` - _Requirements: 2.1_ - [x] 5.3 Update _ENQUEUED_TTL from 14400 to 3600 - File: `services/scheduler/app.py` - Change `_ENQUEUED_TTL = 3600` (1 hour, matching the new recovery cadence) - _Bug_Condition: input.stuckParsedDocCount > 100 AND input.recoveryBatchLimit == 100 AND input.docStaleMinutes >= 240_ - _Expected_Behavior: Recovery processes up to 500 docs per cycle with 30-min threshold_ - _Preservation: Documents < 30 min old left alone for extraction queue consumer_ - _Requirements: 2.1, 3.1, 3.7_ - [x] 6. Fix: Extended Price Fallback - [x] 6.1 Add market_snapshots 24h time-window fallback to create_prediction_snapshot() - File: `services/validation/prediction_snapshot.py` - After positions fallback fails, query: `SELECT close FROM market_snapshots WHERE ticker = $1 AND timestamp >= NOW() - INTERVAL '24 hours' ORDER BY timestamp DESC LIMIT 1` - Add info-level logging when extended fallback is used - _Bug_Condition: input.tickerPrice IS NULL AND input.positionPrice IS NULL AND input.marketSnapshotWithin24h IS NOT NULL_ - _Expected_Behavior: Use market_snapshots bar close price as price_at_prediction_ - _Preservation: Primary fetch_latest_close_price() path unchanged when it returns a valid price_ - _Requirements: 2.2, 3.2_ - [x] 6.2 Create backfill script scripts/backfill_snapshot_prices.py - Query all `prediction_snapshots` where `price_at_prediction IS NULL` - For each, attempt: `market_snapshots` within 24h of `generated_at`, then `positions` table - Update row with found price - Report statistics: found via market_snapshots, found via positions, still NULL - _Requirements: 2.3_ - [x] 7. Fix: Sentiment Z-Score Normalization - [x] 7.1 Add normalize_impact_scores() function - File: `services/aggregation/scoring.py` (new helper function) - Query 7-day mean and stddev per ticker: `SELECT AVG(impact_score) as mean, STDDEV(impact_score) as stddev FROM document_impact_records WHERE ticker = $1 AND created_at >= NOW() - INTERVAL '7 days'` - Formula: `normalized = (raw - mean_7d) / max(stddev_7d, 0.1)` - Fallback: if fewer than 5 records in 7-day window, return raw score unchanged - The 0.1 floor prevents division by near-zero stddev for low-activity tickers - _Requirements: 2.4_ - [x] 7.2 Integrate normalization into aggregation loop - File: `services/aggregation/worker.py` - Before `compute_signal_weight()` call (around line 440), normalize `imp.impact_score` via `normalize_impact_scores()` - Pass normalized value into `compute_signal_weight()` and `WeightedSignal` - _Bug_Condition: input.impactScoreUsedRaw == TRUE AND input.ticker7dStddev > 0_ - _Expected_Behavior: Normalized impact scores with mean ≈ 0, stddev ≈ 1 for active tickers_ - _Preservation: Tickers with balanced distributions continue to produce neutral signals_ - _Requirements: 2.4, 3.3_ - [x] 8. Verify fixes pass all tests - [x] 8.1 Verify bug condition exploration test now passes - **Property 1: Expected Behavior** - Pipeline Health Bugs Resolved - **IMPORTANT**: Re-run the SAME test from task 1 - do NOT write a new test - The test from task 1 encodes the expected behavior for all five bugs - Run `tests/test_pbt_pipeline_health_bug_condition.py` - **EXPECTED OUTCOME**: Test PASSES (confirms bugs are fixed) - _Requirements: 2.1, 2.2, 2.4, 2.6_ - [x] 8.2 Verify preservation tests still pass - **Property 2: Preservation** - Pipeline Behavior Unchanged for Non-Bug Inputs - **IMPORTANT**: Re-run the SAME tests from task 2 - do NOT write new tests - Run `tests/test_pbt_pipeline_health_preservation.py` - **EXPECTED OUTCOME**: Tests PASS (confirms no regressions) - Confirm all preservation properties still hold after fixes - _Requirements: 3.1, 3.2, 3.3, 3.5, 3.6, 3.7_ - [x] 9. Lint and final validation - Run `.venv/bin/ruff check services/` and fix any lint errors - Run `.venv/bin/python -m pytest tests/ -x --tb=short -q` to confirm full test suite passes - Verify Helm template renders correctly with 0 signal-engine replicas - _Requirements: all_ - [x] 10. Checkpoint - Ensure all tests pass - Ensure all tests pass, ask the user if questions arise. ## Task Dependency Graph ```json { "waves": [ {"tasks": ["1", "2"]}, {"tasks": ["3", "4"]}, {"tasks": ["5"]}, {"tasks": ["6"]}, {"tasks": ["7"]}, {"tasks": ["8"]}, {"tasks": ["9"]}, {"tasks": ["10"]} ] } ``` Tasks 3, 4 are independent config changes (no code deps). Task 5 is independent but task 6 builds on the price fallback concept. Task 7 is the most complex (new function + integration). Tasks 8-10 must run after all fixes are applied. ## Notes - Bug 4 (signal engine) is validated by Helm template rendering, not a unit test - The backfill script (6.2) is a one-time operation, not covered by recurring tests - Preservation tests use Hypothesis with `@settings(max_examples=100)` per project conventions - Test files follow `test_pbt_*` naming convention per project standards