Files
stonks-oracle/.kiro/specs/pipeline-health-fixes/design.md
T
Celes Renata ca712ad4a0 fix: pipeline health — stuck docs, price fallback, sentiment normalization, signal-engine scale, quality gate
- Scheduler: lower stale threshold 240→30 min, batch limit 100→500, TTL 14400→3600
- Prediction snapshot: add 24h market_snapshots time-window fallback
- Aggregation: add normalize_impact_scores() z-score normalization
- Helm: signal-engine replicas → 0 (idle when dual pipeline disabled)
- Quality gate: max_snapshot_age_hours 24→48
- Add backfill script for NULL price_at_prediction snapshots
- Add PBT bug condition and preservation tests (14 tests)
2026-07-10 20:16:01 +00:00

321 lines
18 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
# Pipeline Health Fixes — Bugfix Design
## Overview
Five operational bugs degrade stonks-beta pipeline health. This design formalizes the bug conditions, expected fixes, and validation strategy for each:
1. **Stuck Parsed Docs**`recover_stale_documents()` batch limit of 100 is too low for 1,809 stuck documents; increase to 500 and lower the stale threshold to 30 minutes.
2. **Extended Price Fallback** — Prediction snapshots missing prices (26.5%) because the fallback chain stops at `positions`; add a third fallback querying `market_snapshots` within 24h.
3. **Sentiment Z-Score Normalization** — Raw NuExtract3 `impact_score` values produce 64% sell bias; normalize using 7-day rolling z-scores per ticker before signal weighting.
4. **Signal Engine Scale Down** — Idle signal-engine pods consume resources; set replicas to 0 in all Helm values files.
5. **Quality Gate Threshold** — 24h staleness threshold permanently locks quality gate to paper-only; relax to 48h.
## Glossary
- **Bug_Condition (C)**: The specific conditions under which each bug manifests
- **Property (P)**: The desired correct behavior after the fix is applied
- **Preservation**: Existing behavior that must remain unchanged after the fix
- **`recover_stale_documents()`**: Function in `services/scheduler/app.py` that re-enqueues documents stuck in `parsed` status
- **`STALE_PARSED_THRESHOLD_MINUTES`**: Constant (currently 240) controlling how long a document must be stuck before recovery
- **`fetch_latest_close_price()`**: Function in `services/validation/prediction_snapshot.py` that queries `market_snapshots` for the most recent bar
- **`compute_signal_weight()`**: Function in `services/aggregation/scoring.py` that computes combined signal weight from recency, credibility, novelty, confidence, and impact
- **`QualityGateConfig.max_snapshot_age_hours`**: Threshold in `services/trading/model_quality_gate.py` controlling when the quality gate defaults to paper-only
## Bug Details
### Bug Condition
The pipeline health degradation manifests across five independent conditions:
**Formal Specification:**
```
FUNCTION isBugCondition(input)
INPUT: input of type PipelineState
OUTPUT: boolean
-- Bug 1: Parsed docs stuck beyond batch capacity
RETURN (input.stuckParsedDocCount > 100
AND input.recoveryBatchLimit == 100
AND input.docStaleMinutes >= 240)
-- Bug 2: Price fallback chain incomplete
OR (input.tickerPrice IS NULL
AND input.positionPrice IS NULL
AND input.marketSnapshotWithin24h IS NOT NULL)
-- Bug 3: Raw impact scores without normalization
OR (input.impactScoreUsedRaw == TRUE
AND input.ticker7dStddev > 0)
-- Bug 4: Signal engine running idle
OR (input.signalEngineReplicas > 0
AND input.dualPipelineEnabled == FALSE)
-- Bug 5: Quality gate threshold too strict
OR (input.snapshotAgeHours > 24
AND input.snapshotAgeHours <= 48
AND input.maxSnapshotAgeConfig == 24)
END FUNCTION
```
### Examples
- **Bug 1**: 1,809 documents in `parsed` status older than 4 hours. At 100/cycle every 5 minutes, clearing takes 90+ cycles (~7.5h). With 500/batch, it takes 4 cycles (~20 min).
- **Bug 2**: Ticker PLTR has no open position and `fetch_latest_close_price` returns NULL, but `market_snapshots` has a bar from 3 hours ago that could serve as price.
- **Bug 3**: NuExtract3 outputs `impact_score` values clustered around -0.3 to -0.1 for a ticker. Without normalization, `weighted_sentiment_average()` systematically produces negative signals → 64% sell recommendations.
- **Bug 4**: signal-engine pod starts, detects `dual_pipeline_enabled=False`, enters infinite sleep loop consuming 100m CPU request / 128Mi memory request.
- **Bug 5**: Quality gate reads `model_metric_snapshots`, finds the most recent is 26h old (because validation skips NULL-price predictions), fails staleness check, forces paper-only mode permanently.
## Expected Behavior
### Preservation Requirements
**Unchanged Behaviors:**
- Documents entering `parsed` status and processed within the normal threshold window (< 30 min after fix) are left alone for the extraction queue consumer
- Primary price lookup via `fetch_latest_close_price()` (exact time match from `market_snapshots`) continues as the first-choice price source
- Tickers with balanced sentiment distributions continue to produce neutral/mixed recommendations without artificial skew
- Signal-engine's queue-based architecture and configuration loading remain functional when re-enabled with replicas > 0
- Quality gate threshold logic for snapshots younger than 48h and meeting all criteria continues to promote to `live_eligible`
- Outcome evaluator continues to process prediction snapshots with valid (non-NULL) prices normally
- `retry_failed_extractions` task continues on its existing cadence without interference
**Scope:**
All inputs that do NOT match the bug conditions above should be completely unaffected by these fixes. The fixes are additive (new fallback path, wider batch, normalization layer) or config-only (replica count, threshold constant).
## Hypothesized Root Cause
### Bug 1: Stuck Parsed Docs
- **Batch limit too small**: The `LIMIT 100` in the SQL query caps recovery throughput at 100 docs per scheduler cycle (~5 min). When a Redis crash orphans thousands of documents, the recovery rate cannot keep up with the backlog.
- **Threshold too conservative**: `STALE_PARSED_THRESHOLD_MINUTES = 240` (4 hours) means documents must be stuck for 4 hours before recovery kicks in. A 30-minute threshold would catch orphans much faster.
### Bug 2: Incomplete Fallback Chain
- **Missing time-window query**: `fetch_latest_close_price()` only checks `market_snapshots` for an exact timestamp match. When market data ingestion is delayed or the prediction happens outside market hours, no exact match exists.
- **Positions-only fallback**: The positions table fallback only works for tickers with an active position. 26.5% of snapshots are for tickers without positions.
### Bug 3: Uncalibrated Impact Scores
- **No distribution normalization**: NuExtract3 model outputs are passed raw into `compute_signal_weight()` via `impact_score` parameter. The model has a systematic negative bias in its output distribution that is not corrected.
- **Per-ticker variance ignored**: Different tickers receive different volume/types of news, producing different impact_score distributions. A global normalization would be insufficient.
### Bug 4: Idle Signal Engine
- **Replicas set to 1 by default**: `values.yaml` defines `signalEngine.replicas: 1` regardless of whether the dual pipeline feature is enabled. The pod starts, detects the feature is off, and sleeps forever.
### Bug 5: Overly Strict Staleness
- **24h threshold too tight during bootstrapping**: The `max_snapshot_age_hours = 24` default assumes the validation cycle runs frequently. When Bug 2/3 cause most predictions to be skipped, metric snapshots aren't generated, and the 24h window expires.
## Correctness Properties
Property 1: Bug Condition — Stuck Parsed Docs Recovery
_For any_ set of documents stuck in `parsed` status longer than 30 minutes, the fixed `recover_stale_documents()` function SHALL process up to 500 documents per cycle, reducing backlog clearance time by 5x compared to the previous 100-document limit.
**Validates: Requirements 2.1**
Property 2: Bug Condition — Extended Price Fallback
_For any_ prediction snapshot where `fetch_latest_close_price()` returns NULL and the `positions` table has no price, but `market_snapshots` contains a bar for the ticker within 24 hours of the prediction time, the fixed code SHALL use that bar's close price as `price_at_prediction`.
**Validates: Requirements 2.2, 2.3**
Property 3: Bug Condition — Sentiment Z-Score Normalization
_For any_ set of `document_impact_records` for a ticker, the fixed aggregation pipeline SHALL normalize `impact_score` values using the 7-day rolling mean and standard deviation for that ticker before passing them into signal weight computation, preventing systematic model bias.
**Validates: Requirements 2.4**
Property 4: Bug Condition — Signal Engine Scale Down
_For any_ Helm deployment where `dual_pipeline_enabled=False`, the fixed Helm values SHALL specify `signalEngine.replicas: 0`, preventing the pod from being scheduled and consuming resources.
**Validates: Requirements 2.5**
Property 5: Bug Condition — Quality Gate Threshold
_For any_ model metric snapshot that is between 24h and 48h old, the fixed quality gate SHALL NOT reject it as stale, allowing the system to remain in non-paper mode during the bootstrapping period.
**Validates: Requirements 2.6**
Property 6: Preservation — Normal Document Processing
_For any_ document that enters `parsed` status and is processed within 30 minutes, the fixed `recover_stale_documents()` function SHALL NOT interfere with normal extraction queue processing, preserving the existing pipeline flow.
**Validates: Requirements 3.1, 3.7**
Property 7: Preservation — Primary Price Path
_For any_ prediction snapshot where `fetch_latest_close_price()` returns a valid price, the fixed code SHALL use that price directly without invoking any fallback, preserving the primary price lookup behavior.
**Validates: Requirements 3.2**
Property 8: Preservation — Balanced Sentiment
_For any_ ticker with a balanced sentiment distribution (equal bullish/bearish evidence), the z-score normalization SHALL produce values centered around 0, preserving neutral/mixed recommendation output without artificial skew.
**Validates: Requirements 3.3**
Property 9: Preservation — Quality Gate Valid Snapshots
_For any_ model metric snapshot younger than 48h that meets all threshold criteria, the fixed quality gate SHALL continue to promote recommendations to `live_eligible` mode per existing logic.
**Validates: Requirements 3.5, 3.6**
## Fix Implementation
### Changes Required
**Bug 1: Stuck Parsed Docs Recovery**
**File**: `services/scheduler/app.py`
**Function**: `recover_stale_documents()`
**Specific Changes**:
1. **Lower stale threshold**: Change `STALE_PARSED_THRESHOLD_MINUTES` from `240` to `30` — documents stuck longer than 30 minutes are likely orphaned
2. **Increase batch limit**: Change `LIMIT 100` to `LIMIT 500` in the SQL query
3. **Update enqueued TTL**: Change `_ENQUEUED_TTL` from `14400` (4h) to `3600` (1h) to match the new threshold
---
**Bug 2: Extended Price Fallback**
**File**: `services/validation/prediction_snapshot.py`
**Function**: `create_prediction_snapshot()`
**Specific Changes**:
1. **Add market_snapshots time-window fallback**: After the positions fallback fails, query `market_snapshots` for the most recent bar within 24h of the current time for the ticker
2. **SQL query**: `SELECT close FROM market_snapshots WHERE ticker = $1 AND timestamp >= NOW() - INTERVAL '24 hours' ORDER BY timestamp DESC LIMIT 1`
3. **Log the fallback**: Add info-level logging when the extended fallback is used
**New File**: `scripts/backfill_snapshot_prices.py`
**Purpose**: One-time backfill script to populate `price_at_prediction` for existing NULL snapshots using the extended fallback chain.
**Approach**:
1. Query all `prediction_snapshots` where `price_at_prediction IS NULL`
2. For each, attempt: `market_snapshots` within 24h of `generated_at`, then `positions` table
3. Update the row with the found price
4. Report statistics (found via market_snapshots, found via positions, still NULL)
---
**Bug 3: Sentiment Z-Score Normalization**
**File**: `services/aggregation/worker.py` (or new helper in `services/aggregation/scoring.py`)
**Function**: New function `normalize_impact_scores()` called before `compute_signal_weight()`
**Specific Changes**:
1. **Add normalization function**: Compute 7-day rolling mean and stddev of `impact_score` per ticker from `document_impact_records`
2. **Formula**: `normalized = (raw - mean_7d) / max(stddev_7d, 0.1)` — the 0.1 floor prevents division by near-zero stddev for low-activity tickers
3. **Integration point**: In the aggregation loop (around line 440 of worker.py), normalize `imp.impact_score` before passing to `compute_signal_weight()` and `WeightedSignal`
4. **Fallback**: If fewer than 5 records exist in the 7-day window, use the raw score (insufficient data for meaningful normalization)
5. **Query**: `SELECT AVG(impact_score) as mean, STDDEV(impact_score) as stddev FROM document_impact_records WHERE ticker = $1 AND created_at >= NOW() - INTERVAL '7 days'`
---
**Bug 4: Signal Engine Scale Down**
**Files**: `infra/helm/stonks-oracle/values.yaml`, `values-beta.yaml`, `values-paper.yaml`
**Specific Changes**:
1. **values.yaml**: Change `signalEngine.replicas` from `1` to `0`
2. **values-beta.yaml**: Add `signalEngine.replicas: 0` under `services:`
3. **values-paper.yaml**: Add `signalEngine.replicas: 0` under `services:`
---
**Bug 5: Quality Gate Threshold**
**File**: `services/trading/model_quality_gate.py`
**Class**: `QualityGateConfig`
**Specific Changes**:
1. **Change default**: `max_snapshot_age_hours: int = 48` (was 24)
## Testing Strategy
### Validation Approach
The testing strategy follows a two-phase approach: first, surface counterexamples that demonstrate the bug on unfixed code, then verify the fix works correctly and preserves existing behavior.
### Exploratory Bug Condition Checking
**Goal**: Surface counterexamples that demonstrate the bugs BEFORE implementing the fixes. Confirm or refute the root cause analysis.
**Test Plan**: Write tests that exercise each bug condition on the unfixed code to observe failures.
**Test Cases**:
1. **Batch Overflow Test**: Create 600 documents in `parsed` status older than threshold, run `recover_stale_documents()`, assert only 100 are processed (will demonstrate Bug 1)
2. **Price Fallback Gap Test**: Call `create_prediction_snapshot()` for a ticker with no position and no exact market_snapshots match, assert `price_at_prediction` is NULL (will demonstrate Bug 2)
3. **Sentiment Bias Test**: Generate 50 impact records with systematic negative bias (mean=-0.3, stddev=0.1), compute weighted signals, assert directional signal is negative (will demonstrate Bug 3)
4. **Quality Gate Staleness Test**: Set most recent metric snapshot to 26h ago, evaluate quality gate, assert it fails (will demonstrate Bug 5)
**Expected Counterexamples**:
- Bug 1: Only 100 of 600 documents recovered per cycle
- Bug 2: `price_at_prediction` stored as NULL despite market data existing within 24h
- Bug 3: Weighted sentiment average heavily negative despite mixed underlying events
- Bug 5: Quality gate returns `passed=False` with reason containing "stale"
### Fix Checking
**Goal**: Verify that for all inputs where the bug condition holds, the fixed function produces the expected behavior.
**Pseudocode:**
```
FOR ALL input WHERE isBugCondition(input) DO
result := fixedFunction(input)
ASSERT expectedBehavior(result)
END FOR
```
**Per-bug fix checks:**
- Bug 1: `recover_stale_documents()` processes up to 500 docs with 30-min threshold
- Bug 2: Extended fallback returns a price when `market_snapshots` has data within 24h
- Bug 3: Normalized impact scores have mean ≈ 0 and stddev ≈ 1 for active tickers
- Bug 4: `kubectl get pods` shows 0 signal-engine pods
- Bug 5: Quality gate passes for snapshots 2448h old that meet metric thresholds
### Preservation Checking
**Goal**: Verify that for all inputs where the bug condition does NOT hold, the fixed function produces the same result as the original function.
**Pseudocode:**
```
FOR ALL input WHERE NOT isBugCondition(input) DO
ASSERT originalFunction(input) = fixedFunction(input)
END FOR
```
**Testing Approach**: Property-based testing is recommended for preservation checking because:
- It generates many test cases automatically across the input domain
- It catches edge cases that manual unit tests might miss
- It provides strong guarantees that behavior is unchanged for all non-buggy inputs
**Test Plan**: Observe behavior on UNFIXED code first for normal inputs, then write property-based tests capturing that behavior.
**Test Cases**:
1. **Normal Doc Processing Preservation**: Documents < 30 min old are never touched by recovery
2. **Primary Price Preservation**: When `fetch_latest_close_price()` succeeds, no fallback is invoked
3. **Balanced Sentiment Preservation**: Tickers with symmetric impact_score distributions produce neutral signals after normalization
4. **Quality Gate Normal Preservation**: Snapshots < 48h old and meeting thresholds still pass
5. **Failed Extraction Preservation**: `retry_failed_extractions()` behavior unchanged
### Unit Tests
- Test `recover_stale_documents()` with various document counts (0, 50, 500, 1000)
- Test extended price fallback with market_snapshots at various time offsets (1h, 12h, 23h, 25h)
- Test z-score normalization with known distributions (mean=0, mean=-0.5, stddev=0, stddev=0.05)
- Test quality gate with snapshot ages at boundary (23h, 24h, 47h, 48h, 49h)
- Test backfill script with mixed NULL/non-NULL snapshots
### Property-Based Tests
- Generate random document ages and counts, verify recovery processes correct subset (> 30 min old, up to 500)
- Generate random ticker price scenarios, verify fallback chain ordering is preserved (primary → positions → market_snapshots_24h → NULL)
- Generate random impact_score distributions per ticker, verify normalized output has bounded variance and zero-centered mean
- Generate random snapshot ages, verify quality gate accepts [0, 48h) and rejects [48h, ∞)
### Integration Tests
- End-to-end: create documents in `parsed` status, run scheduler cycle, verify extraction queue populated
- End-to-end: create prediction snapshot for ticker without position, verify price populated from market_snapshots
- End-to-end: run full aggregation cycle with biased NuExtract3 outputs, verify recommendation direction is not systematically biased
- Helm template render: verify signal-engine deployment has 0 replicas in all value files