fix: pipeline health — stuck docs, price fallback, sentiment normalization, signal-engine scale, quality gate

- Scheduler: lower stale threshold 240→30 min, batch limit 100→500, TTL 14400→3600
- Prediction snapshot: add 24h market_snapshots time-window fallback
- Aggregation: add normalize_impact_scores() z-score normalization
- Helm: signal-engine replicas → 0 (idle when dual pipeline disabled)
- Quality gate: max_snapshot_age_hours 24→48
- Add backfill script for NULL price_at_prediction snapshots
- Add PBT bug condition and preservation tests (14 tests)
This commit is contained in:
Celes Renata
2026-07-10 20:16:01 +00:00
parent a4f51c00e1
commit ca712ad4a0
15 changed files with 1815 additions and 9 deletions
@@ -0,0 +1 @@
{"specId": "f5d99301-94ef-4dc2-8ba4-ccefeee7ecba", "workflowType": "requirements-first", "specType": "bugfix"}
@@ -0,0 +1,51 @@
# Bugfix Requirements Document
## Introduction
Five operational bugs in the stonks-beta deployment degrade pipeline health: 1,809 documents stuck in `parsed` status due to recovery batch limits, 26.5% of prediction snapshots missing prices due to incomplete fallback chains, 64% sell bias from uncalibrated NuExtract3 sentiment outputs, idle signal-engine consuming resources while doing nothing, and a quality gate stuck in paper-only mode due to an overly strict staleness threshold interacting with the NULL price problem.
## Bug Analysis
### Current Behavior (Defect)
1.1 WHEN the `recover_stale_documents` task runs with 1,809+ documents stuck in `parsed` status THEN the system only processes 100 per cycle (every ~5 minutes), requiring 90+ cycles (~7.5 hours) to clear the backlog while new documents may continue accumulating
1.2 WHEN a prediction snapshot is created for a ticker without an open position AND without recent market_snapshots data THEN the system stores NULL in `price_at_prediction` because the fallback chain stops at the positions table (26.5% of snapshots affected — 33,324 of 125,590)
1.3 WHEN the outcome evaluator encounters a prediction snapshot with NULL `price_at_prediction` THEN the system skips the snapshot entirely, creating a validation blind spot where 26.5% of predictions are never evaluated
1.4 WHEN the aggregation pipeline processes NuExtract3 extraction outputs THEN the system passes raw `impact_score` and `sentiment` values directly into signal weighting without any distribution normalization, resulting in systematic negative bias producing 64% sell / 23% watch / 12% buy recommendations
1.5 WHEN the signal-engine pod starts with `dual_pipeline_enabled=False` THEN the system enters an infinite sleep loop consuming CPU (100m request / 500m limit) and memory (128Mi request / 256Mi limit) while producing zero signal evaluations
1.6 WHEN the quality gate checks `model_metric_snapshots` freshness with a 24-hour staleness threshold AND the validation cycle skips all predictions due to NULL prices (Bug 1.2/1.3) THEN the system permanently defaults to paper-only mode because no fresh metric snapshots are ever generated
### Expected Behavior (Correct)
2.1 WHEN the scheduler detects more than 100 documents stuck in `parsed` status older than the threshold THEN the system SHALL increase the batch limit for recovery processing (up to 500 per cycle) and provide a one-time management command to bulk-recover the existing backlog without waiting for periodic sweeps
2.2 WHEN a prediction snapshot is created and no price is available from market_snapshots (exact time) or positions table THEN the system SHALL query `market_snapshots` with a wider time window (last 24 hours of bar data for the ticker) as an additional fallback before accepting NULL
2.3 WHEN backfilling existing prediction snapshots with NULL `price_at_prediction` THEN the system SHALL use the extended fallback chain (market_snapshots within 24h of `generated_at`, then positions) to populate prices retroactively via a migration script
2.4 WHEN the aggregation pipeline computes signal weights from impact records THEN the system SHALL apply z-score normalization to `impact_score` values relative to the rolling 7-day distribution of impact records for the same ticker, preventing systematic model bias from dominating the directional signal
2.5 WHEN the signal-engine deployment is not ready for production use (`dual_pipeline_enabled=False`) THEN the system SHALL be scaled to 0 replicas in the Helm values files (beta, paper, live) to eliminate wasted CPU, memory, and any GPU time-slice allocations
2.6 WHEN the quality gate evaluates metric snapshot freshness during the bootstrapping period THEN the system SHALL use a 48-hour staleness threshold (instead of 24h) to tolerate gaps while the validation cycle ramps up after Bug 1.2/1.3 are fixed
### Unchanged Behavior (Regression Prevention)
3.1 WHEN documents enter `parsed` status and are processed within the normal threshold window (< 240 minutes) THEN the system SHALL CONTINUE TO leave them for the extraction queue consumer without interference from the recovery task
3.2 WHEN a prediction snapshot is created and market_snapshots contains a recent bar for the ticker THEN the system SHALL CONTINUE TO use the primary `market_snapshots` close price without invoking any fallback
3.3 WHEN the aggregation pipeline processes tickers with balanced sentiment distributions (equal bullish/bearish evidence) THEN the system SHALL CONTINUE TO produce neutral/mixed recommendations without artificial skew from the normalization step
3.4 WHEN the signal-engine is re-enabled in the future (dual_pipeline_enabled=True with replicas > 0) THEN the system SHALL CONTINUE TO function correctly with its existing queue-based architecture and configuration loading
3.5 WHEN the quality gate evaluates a metric snapshot that is less than 48 hours old and meets all threshold criteria THEN the system SHALL CONTINUE TO promote recommendations to live_eligible mode per existing threshold logic
3.6 WHEN the outcome evaluator processes prediction snapshots with valid (non-NULL) prices THEN the system SHALL CONTINUE TO evaluate them normally and produce prediction_outcomes records
3.7 WHEN the `retry_failed_extractions` task handles documents in `extraction_failed` status THEN the system SHALL CONTINUE TO process them on the existing cadence and logic without interference from the parsed-document recovery changes
+320
View File
@@ -0,0 +1,320 @@
# Pipeline Health Fixes — Bugfix Design
## Overview
Five operational bugs degrade stonks-beta pipeline health. This design formalizes the bug conditions, expected fixes, and validation strategy for each:
1. **Stuck Parsed Docs**`recover_stale_documents()` batch limit of 100 is too low for 1,809 stuck documents; increase to 500 and lower the stale threshold to 30 minutes.
2. **Extended Price Fallback** — Prediction snapshots missing prices (26.5%) because the fallback chain stops at `positions`; add a third fallback querying `market_snapshots` within 24h.
3. **Sentiment Z-Score Normalization** — Raw NuExtract3 `impact_score` values produce 64% sell bias; normalize using 7-day rolling z-scores per ticker before signal weighting.
4. **Signal Engine Scale Down** — Idle signal-engine pods consume resources; set replicas to 0 in all Helm values files.
5. **Quality Gate Threshold** — 24h staleness threshold permanently locks quality gate to paper-only; relax to 48h.
## Glossary
- **Bug_Condition (C)**: The specific conditions under which each bug manifests
- **Property (P)**: The desired correct behavior after the fix is applied
- **Preservation**: Existing behavior that must remain unchanged after the fix
- **`recover_stale_documents()`**: Function in `services/scheduler/app.py` that re-enqueues documents stuck in `parsed` status
- **`STALE_PARSED_THRESHOLD_MINUTES`**: Constant (currently 240) controlling how long a document must be stuck before recovery
- **`fetch_latest_close_price()`**: Function in `services/validation/prediction_snapshot.py` that queries `market_snapshots` for the most recent bar
- **`compute_signal_weight()`**: Function in `services/aggregation/scoring.py` that computes combined signal weight from recency, credibility, novelty, confidence, and impact
- **`QualityGateConfig.max_snapshot_age_hours`**: Threshold in `services/trading/model_quality_gate.py` controlling when the quality gate defaults to paper-only
## Bug Details
### Bug Condition
The pipeline health degradation manifests across five independent conditions:
**Formal Specification:**
```
FUNCTION isBugCondition(input)
INPUT: input of type PipelineState
OUTPUT: boolean
-- Bug 1: Parsed docs stuck beyond batch capacity
RETURN (input.stuckParsedDocCount > 100
AND input.recoveryBatchLimit == 100
AND input.docStaleMinutes >= 240)
-- Bug 2: Price fallback chain incomplete
OR (input.tickerPrice IS NULL
AND input.positionPrice IS NULL
AND input.marketSnapshotWithin24h IS NOT NULL)
-- Bug 3: Raw impact scores without normalization
OR (input.impactScoreUsedRaw == TRUE
AND input.ticker7dStddev > 0)
-- Bug 4: Signal engine running idle
OR (input.signalEngineReplicas > 0
AND input.dualPipelineEnabled == FALSE)
-- Bug 5: Quality gate threshold too strict
OR (input.snapshotAgeHours > 24
AND input.snapshotAgeHours <= 48
AND input.maxSnapshotAgeConfig == 24)
END FUNCTION
```
### Examples
- **Bug 1**: 1,809 documents in `parsed` status older than 4 hours. At 100/cycle every 5 minutes, clearing takes 90+ cycles (~7.5h). With 500/batch, it takes 4 cycles (~20 min).
- **Bug 2**: Ticker PLTR has no open position and `fetch_latest_close_price` returns NULL, but `market_snapshots` has a bar from 3 hours ago that could serve as price.
- **Bug 3**: NuExtract3 outputs `impact_score` values clustered around -0.3 to -0.1 for a ticker. Without normalization, `weighted_sentiment_average()` systematically produces negative signals → 64% sell recommendations.
- **Bug 4**: signal-engine pod starts, detects `dual_pipeline_enabled=False`, enters infinite sleep loop consuming 100m CPU request / 128Mi memory request.
- **Bug 5**: Quality gate reads `model_metric_snapshots`, finds the most recent is 26h old (because validation skips NULL-price predictions), fails staleness check, forces paper-only mode permanently.
## Expected Behavior
### Preservation Requirements
**Unchanged Behaviors:**
- Documents entering `parsed` status and processed within the normal threshold window (< 30 min after fix) are left alone for the extraction queue consumer
- Primary price lookup via `fetch_latest_close_price()` (exact time match from `market_snapshots`) continues as the first-choice price source
- Tickers with balanced sentiment distributions continue to produce neutral/mixed recommendations without artificial skew
- Signal-engine's queue-based architecture and configuration loading remain functional when re-enabled with replicas > 0
- Quality gate threshold logic for snapshots younger than 48h and meeting all criteria continues to promote to `live_eligible`
- Outcome evaluator continues to process prediction snapshots with valid (non-NULL) prices normally
- `retry_failed_extractions` task continues on its existing cadence without interference
**Scope:**
All inputs that do NOT match the bug conditions above should be completely unaffected by these fixes. The fixes are additive (new fallback path, wider batch, normalization layer) or config-only (replica count, threshold constant).
## Hypothesized Root Cause
### Bug 1: Stuck Parsed Docs
- **Batch limit too small**: The `LIMIT 100` in the SQL query caps recovery throughput at 100 docs per scheduler cycle (~5 min). When a Redis crash orphans thousands of documents, the recovery rate cannot keep up with the backlog.
- **Threshold too conservative**: `STALE_PARSED_THRESHOLD_MINUTES = 240` (4 hours) means documents must be stuck for 4 hours before recovery kicks in. A 30-minute threshold would catch orphans much faster.
### Bug 2: Incomplete Fallback Chain
- **Missing time-window query**: `fetch_latest_close_price()` only checks `market_snapshots` for an exact timestamp match. When market data ingestion is delayed or the prediction happens outside market hours, no exact match exists.
- **Positions-only fallback**: The positions table fallback only works for tickers with an active position. 26.5% of snapshots are for tickers without positions.
### Bug 3: Uncalibrated Impact Scores
- **No distribution normalization**: NuExtract3 model outputs are passed raw into `compute_signal_weight()` via `impact_score` parameter. The model has a systematic negative bias in its output distribution that is not corrected.
- **Per-ticker variance ignored**: Different tickers receive different volume/types of news, producing different impact_score distributions. A global normalization would be insufficient.
### Bug 4: Idle Signal Engine
- **Replicas set to 1 by default**: `values.yaml` defines `signalEngine.replicas: 1` regardless of whether the dual pipeline feature is enabled. The pod starts, detects the feature is off, and sleeps forever.
### Bug 5: Overly Strict Staleness
- **24h threshold too tight during bootstrapping**: The `max_snapshot_age_hours = 24` default assumes the validation cycle runs frequently. When Bug 2/3 cause most predictions to be skipped, metric snapshots aren't generated, and the 24h window expires.
## Correctness Properties
Property 1: Bug Condition — Stuck Parsed Docs Recovery
_For any_ set of documents stuck in `parsed` status longer than 30 minutes, the fixed `recover_stale_documents()` function SHALL process up to 500 documents per cycle, reducing backlog clearance time by 5x compared to the previous 100-document limit.
**Validates: Requirements 2.1**
Property 2: Bug Condition — Extended Price Fallback
_For any_ prediction snapshot where `fetch_latest_close_price()` returns NULL and the `positions` table has no price, but `market_snapshots` contains a bar for the ticker within 24 hours of the prediction time, the fixed code SHALL use that bar's close price as `price_at_prediction`.
**Validates: Requirements 2.2, 2.3**
Property 3: Bug Condition — Sentiment Z-Score Normalization
_For any_ set of `document_impact_records` for a ticker, the fixed aggregation pipeline SHALL normalize `impact_score` values using the 7-day rolling mean and standard deviation for that ticker before passing them into signal weight computation, preventing systematic model bias.
**Validates: Requirements 2.4**
Property 4: Bug Condition — Signal Engine Scale Down
_For any_ Helm deployment where `dual_pipeline_enabled=False`, the fixed Helm values SHALL specify `signalEngine.replicas: 0`, preventing the pod from being scheduled and consuming resources.
**Validates: Requirements 2.5**
Property 5: Bug Condition — Quality Gate Threshold
_For any_ model metric snapshot that is between 24h and 48h old, the fixed quality gate SHALL NOT reject it as stale, allowing the system to remain in non-paper mode during the bootstrapping period.
**Validates: Requirements 2.6**
Property 6: Preservation — Normal Document Processing
_For any_ document that enters `parsed` status and is processed within 30 minutes, the fixed `recover_stale_documents()` function SHALL NOT interfere with normal extraction queue processing, preserving the existing pipeline flow.
**Validates: Requirements 3.1, 3.7**
Property 7: Preservation — Primary Price Path
_For any_ prediction snapshot where `fetch_latest_close_price()` returns a valid price, the fixed code SHALL use that price directly without invoking any fallback, preserving the primary price lookup behavior.
**Validates: Requirements 3.2**
Property 8: Preservation — Balanced Sentiment
_For any_ ticker with a balanced sentiment distribution (equal bullish/bearish evidence), the z-score normalization SHALL produce values centered around 0, preserving neutral/mixed recommendation output without artificial skew.
**Validates: Requirements 3.3**
Property 9: Preservation — Quality Gate Valid Snapshots
_For any_ model metric snapshot younger than 48h that meets all threshold criteria, the fixed quality gate SHALL continue to promote recommendations to `live_eligible` mode per existing logic.
**Validates: Requirements 3.5, 3.6**
## Fix Implementation
### Changes Required
**Bug 1: Stuck Parsed Docs Recovery**
**File**: `services/scheduler/app.py`
**Function**: `recover_stale_documents()`
**Specific Changes**:
1. **Lower stale threshold**: Change `STALE_PARSED_THRESHOLD_MINUTES` from `240` to `30` — documents stuck longer than 30 minutes are likely orphaned
2. **Increase batch limit**: Change `LIMIT 100` to `LIMIT 500` in the SQL query
3. **Update enqueued TTL**: Change `_ENQUEUED_TTL` from `14400` (4h) to `3600` (1h) to match the new threshold
---
**Bug 2: Extended Price Fallback**
**File**: `services/validation/prediction_snapshot.py`
**Function**: `create_prediction_snapshot()`
**Specific Changes**:
1. **Add market_snapshots time-window fallback**: After the positions fallback fails, query `market_snapshots` for the most recent bar within 24h of the current time for the ticker
2. **SQL query**: `SELECT close FROM market_snapshots WHERE ticker = $1 AND timestamp >= NOW() - INTERVAL '24 hours' ORDER BY timestamp DESC LIMIT 1`
3. **Log the fallback**: Add info-level logging when the extended fallback is used
**New File**: `scripts/backfill_snapshot_prices.py`
**Purpose**: One-time backfill script to populate `price_at_prediction` for existing NULL snapshots using the extended fallback chain.
**Approach**:
1. Query all `prediction_snapshots` where `price_at_prediction IS NULL`
2. For each, attempt: `market_snapshots` within 24h of `generated_at`, then `positions` table
3. Update the row with the found price
4. Report statistics (found via market_snapshots, found via positions, still NULL)
---
**Bug 3: Sentiment Z-Score Normalization**
**File**: `services/aggregation/worker.py` (or new helper in `services/aggregation/scoring.py`)
**Function**: New function `normalize_impact_scores()` called before `compute_signal_weight()`
**Specific Changes**:
1. **Add normalization function**: Compute 7-day rolling mean and stddev of `impact_score` per ticker from `document_impact_records`
2. **Formula**: `normalized = (raw - mean_7d) / max(stddev_7d, 0.1)` — the 0.1 floor prevents division by near-zero stddev for low-activity tickers
3. **Integration point**: In the aggregation loop (around line 440 of worker.py), normalize `imp.impact_score` before passing to `compute_signal_weight()` and `WeightedSignal`
4. **Fallback**: If fewer than 5 records exist in the 7-day window, use the raw score (insufficient data for meaningful normalization)
5. **Query**: `SELECT AVG(impact_score) as mean, STDDEV(impact_score) as stddev FROM document_impact_records WHERE ticker = $1 AND created_at >= NOW() - INTERVAL '7 days'`
---
**Bug 4: Signal Engine Scale Down**
**Files**: `infra/helm/stonks-oracle/values.yaml`, `values-beta.yaml`, `values-paper.yaml`
**Specific Changes**:
1. **values.yaml**: Change `signalEngine.replicas` from `1` to `0`
2. **values-beta.yaml**: Add `signalEngine.replicas: 0` under `services:`
3. **values-paper.yaml**: Add `signalEngine.replicas: 0` under `services:`
---
**Bug 5: Quality Gate Threshold**
**File**: `services/trading/model_quality_gate.py`
**Class**: `QualityGateConfig`
**Specific Changes**:
1. **Change default**: `max_snapshot_age_hours: int = 48` (was 24)
## Testing Strategy
### Validation Approach
The testing strategy follows a two-phase approach: first, surface counterexamples that demonstrate the bug on unfixed code, then verify the fix works correctly and preserves existing behavior.
### Exploratory Bug Condition Checking
**Goal**: Surface counterexamples that demonstrate the bugs BEFORE implementing the fixes. Confirm or refute the root cause analysis.
**Test Plan**: Write tests that exercise each bug condition on the unfixed code to observe failures.
**Test Cases**:
1. **Batch Overflow Test**: Create 600 documents in `parsed` status older than threshold, run `recover_stale_documents()`, assert only 100 are processed (will demonstrate Bug 1)
2. **Price Fallback Gap Test**: Call `create_prediction_snapshot()` for a ticker with no position and no exact market_snapshots match, assert `price_at_prediction` is NULL (will demonstrate Bug 2)
3. **Sentiment Bias Test**: Generate 50 impact records with systematic negative bias (mean=-0.3, stddev=0.1), compute weighted signals, assert directional signal is negative (will demonstrate Bug 3)
4. **Quality Gate Staleness Test**: Set most recent metric snapshot to 26h ago, evaluate quality gate, assert it fails (will demonstrate Bug 5)
**Expected Counterexamples**:
- Bug 1: Only 100 of 600 documents recovered per cycle
- Bug 2: `price_at_prediction` stored as NULL despite market data existing within 24h
- Bug 3: Weighted sentiment average heavily negative despite mixed underlying events
- Bug 5: Quality gate returns `passed=False` with reason containing "stale"
### Fix Checking
**Goal**: Verify that for all inputs where the bug condition holds, the fixed function produces the expected behavior.
**Pseudocode:**
```
FOR ALL input WHERE isBugCondition(input) DO
result := fixedFunction(input)
ASSERT expectedBehavior(result)
END FOR
```
**Per-bug fix checks:**
- Bug 1: `recover_stale_documents()` processes up to 500 docs with 30-min threshold
- Bug 2: Extended fallback returns a price when `market_snapshots` has data within 24h
- Bug 3: Normalized impact scores have mean ≈ 0 and stddev ≈ 1 for active tickers
- Bug 4: `kubectl get pods` shows 0 signal-engine pods
- Bug 5: Quality gate passes for snapshots 2448h old that meet metric thresholds
### Preservation Checking
**Goal**: Verify that for all inputs where the bug condition does NOT hold, the fixed function produces the same result as the original function.
**Pseudocode:**
```
FOR ALL input WHERE NOT isBugCondition(input) DO
ASSERT originalFunction(input) = fixedFunction(input)
END FOR
```
**Testing Approach**: Property-based testing is recommended for preservation checking because:
- It generates many test cases automatically across the input domain
- It catches edge cases that manual unit tests might miss
- It provides strong guarantees that behavior is unchanged for all non-buggy inputs
**Test Plan**: Observe behavior on UNFIXED code first for normal inputs, then write property-based tests capturing that behavior.
**Test Cases**:
1. **Normal Doc Processing Preservation**: Documents < 30 min old are never touched by recovery
2. **Primary Price Preservation**: When `fetch_latest_close_price()` succeeds, no fallback is invoked
3. **Balanced Sentiment Preservation**: Tickers with symmetric impact_score distributions produce neutral signals after normalization
4. **Quality Gate Normal Preservation**: Snapshots < 48h old and meeting thresholds still pass
5. **Failed Extraction Preservation**: `retry_failed_extractions()` behavior unchanged
### Unit Tests
- Test `recover_stale_documents()` with various document counts (0, 50, 500, 1000)
- Test extended price fallback with market_snapshots at various time offsets (1h, 12h, 23h, 25h)
- Test z-score normalization with known distributions (mean=0, mean=-0.5, stddev=0, stddev=0.05)
- Test quality gate with snapshot ages at boundary (23h, 24h, 47h, 48h, 49h)
- Test backfill script with mixed NULL/non-NULL snapshots
### Property-Based Tests
- Generate random document ages and counts, verify recovery processes correct subset (> 30 min old, up to 500)
- Generate random ticker price scenarios, verify fallback chain ordering is preserved (primary → positions → market_snapshots_24h → NULL)
- Generate random impact_score distributions per ticker, verify normalized output has bounded variance and zero-centered mean
- Generate random snapshot ages, verify quality gate accepts [0, 48h) and rejects [48h, ∞)
### Integration Tests
- End-to-end: create documents in `parsed` status, run scheduler cycle, verify extraction queue populated
- End-to-end: create prediction snapshot for ticker without position, verify price populated from market_snapshots
- End-to-end: run full aggregation cycle with biased NuExtract3 outputs, verify recommendation direction is not systematically biased
- Helm template render: verify signal-engine deployment has 0 replicas in all value files
+174
View File
@@ -0,0 +1,174 @@
# Implementation Plan
## Overview
Bugfix implementation for five pipeline health issues: stuck parsed docs, missing price fallback, uncalibrated sentiment scores, idle signal-engine pods, and overly strict quality gate threshold. Tasks follow the exploratory bugfix workflow: explore bugs via tests, preserve existing behavior, implement fixes, validate.
## Tasks
- [x] 1. Write bug condition exploration test
- **Property 1: Bug Condition** - Pipeline Health Degradation
- **CRITICAL**: This test MUST FAIL on unfixed code - failure confirms the bugs exist
- **DO NOT attempt to fix the test or the code when it fails**
- **NOTE**: This test encodes the expected behavior - it will validate the fix when it passes after implementation
- **GOAL**: Surface counterexamples that demonstrate all five bugs exist
- **Scoped PBT Approach**: Scope properties to the concrete failing cases for each bug condition
- Test file: `tests/test_pbt_pipeline_health_bug_condition.py`
- **Bug 1 - Batch Overflow**: Create 600 documents in `parsed` status older than 30 min, run `recover_stale_documents()`, assert up to 500 are recovered per cycle (will FAIL on unfixed code which caps at 100)
- **Bug 2 - Price Fallback Gap**: Call `create_prediction_snapshot()` for ticker with no position and no exact market_snapshots match but data within 24h exists, assert `price_at_prediction` is NOT NULL (will FAIL on unfixed code which returns NULL)
- **Bug 3 - Sentiment Bias**: Generate 50 impact records with systematic negative bias (mean=-0.3, stddev=0.1), compute weighted signals via aggregation, assert normalized output is zero-centered (will FAIL on unfixed code which passes raw scores)
- **Bug 5 - Quality Gate Staleness**: Set most recent metric snapshot to 26h ago, evaluate quality gate, assert it passes (will FAIL on unfixed code which rejects at 24h)
- Run tests on UNFIXED code
- **EXPECTED OUTCOME**: Tests FAIL (this is correct - it proves the bugs exist)
- Document counterexamples: batch capped at 100, price stored as NULL, sentiment heavily negative, quality gate returns `passed=False`
- Mark task complete when tests are written, run, and failures are documented
- _Requirements: 1.1, 1.2, 1.3, 1.4, 1.6_
- [x] 2. Write preservation property tests (BEFORE implementing fix)
- **Property 2: Preservation** - Pipeline Behavior Unchanged for Non-Bug Inputs
- **IMPORTANT**: Follow observation-first methodology
- Test file: `tests/test_pbt_pipeline_health_preservation.py`
- **Normal Doc Processing**: Observe that documents < 30 min old are never touched by `recover_stale_documents()` on unfixed code. Write property: for all documents with age < 30 min, recovery task does NOT enqueue them.
- **Primary Price Path**: Observe that when `fetch_latest_close_price()` returns a valid price, no fallback is invoked. Write property: for all tickers where primary price exists, result equals primary price.
- **Balanced Sentiment**: Observe that tickers with symmetric impact_score distributions (mean ≈ 0) produce neutral signals. Write property: for all impact_score sets with mean ≈ 0, normalized output remains centered around 0.
- **Quality Gate Normal**: Observe that snapshots < 48h old meeting thresholds pass the quality gate. Write property: for all snapshot ages in [0, 48h) meeting metric criteria, quality gate returns `passed=True`.
- **Failed Extraction Independence**: Observe `retry_failed_extractions()` behavior is unaffected. Write property: for all documents in `extraction_failed` status, retry logic unchanged.
- Run tests on UNFIXED code
- **EXPECTED OUTCOME**: Tests PASS (this confirms baseline behavior to preserve)
- Mark task complete when tests are written, run, and passing on unfixed code
- _Requirements: 3.1, 3.2, 3.3, 3.5, 3.6, 3.7_
- [x] 3. Fix: Signal Engine Scale Down (Helm values)
- [x] 3.1 Set signal-engine replicas to 0 in all Helm values files
- In `infra/helm/stonks-oracle/values.yaml`: change `signalEngine.replicas` from `1` to `0`
- In `infra/helm/stonks-oracle/values-beta.yaml`: add/set `signalEngine.replicas: 0` under `services:`
- In `infra/helm/stonks-oracle/values-paper.yaml`: add/set `signalEngine.replicas: 0` under `services:`
- _Bug_Condition: input.signalEngineReplicas > 0 AND input.dualPipelineEnabled == FALSE_
- _Expected_Behavior: signalEngine.replicas == 0 when dual pipeline disabled_
- _Preservation: Signal-engine architecture remains functional when re-enabled with replicas > 0_
- _Requirements: 2.5, 3.4_
- [x] 4. Fix: Quality Gate Threshold Relaxation
- [x] 4.1 Change max_snapshot_age_hours default from 24 to 48
- File: `services/trading/model_quality_gate.py`
- In `QualityGateConfig` class, change `max_snapshot_age_hours: int = 24` to `max_snapshot_age_hours: int = 48`
- _Bug_Condition: input.snapshotAgeHours > 24 AND input.snapshotAgeHours <= 48 AND input.maxSnapshotAgeConfig == 24_
- _Expected_Behavior: Quality gate accepts snapshots up to 48h old_
- _Preservation: Snapshots < 48h meeting criteria continue to promote to live_eligible_
- _Requirements: 2.6, 3.5_
- [x] 5. Fix: Stuck Parsed Docs Recovery
- [x] 5.1 Lower STALE_PARSED_THRESHOLD_MINUTES from 240 to 30
- File: `services/scheduler/app.py`
- Change constant: `STALE_PARSED_THRESHOLD_MINUTES = 30`
- Documents stuck longer than 30 minutes are likely orphaned
- _Requirements: 2.1_
- [x] 5.2 Increase recovery batch LIMIT from 100 to 500
- File: `services/scheduler/app.py`
- In `recover_stale_documents()` SQL query, change `LIMIT 100` to `LIMIT 500`
- _Requirements: 2.1_
- [x] 5.3 Update _ENQUEUED_TTL from 14400 to 3600
- File: `services/scheduler/app.py`
- Change `_ENQUEUED_TTL = 3600` (1 hour, matching the new recovery cadence)
- _Bug_Condition: input.stuckParsedDocCount > 100 AND input.recoveryBatchLimit == 100 AND input.docStaleMinutes >= 240_
- _Expected_Behavior: Recovery processes up to 500 docs per cycle with 30-min threshold_
- _Preservation: Documents < 30 min old left alone for extraction queue consumer_
- _Requirements: 2.1, 3.1, 3.7_
- [x] 6. Fix: Extended Price Fallback
- [x] 6.1 Add market_snapshots 24h time-window fallback to create_prediction_snapshot()
- File: `services/validation/prediction_snapshot.py`
- After positions fallback fails, query: `SELECT close FROM market_snapshots WHERE ticker = $1 AND timestamp >= NOW() - INTERVAL '24 hours' ORDER BY timestamp DESC LIMIT 1`
- Add info-level logging when extended fallback is used
- _Bug_Condition: input.tickerPrice IS NULL AND input.positionPrice IS NULL AND input.marketSnapshotWithin24h IS NOT NULL_
- _Expected_Behavior: Use market_snapshots bar close price as price_at_prediction_
- _Preservation: Primary fetch_latest_close_price() path unchanged when it returns a valid price_
- _Requirements: 2.2, 3.2_
- [x] 6.2 Create backfill script scripts/backfill_snapshot_prices.py
- Query all `prediction_snapshots` where `price_at_prediction IS NULL`
- For each, attempt: `market_snapshots` within 24h of `generated_at`, then `positions` table
- Update row with found price
- Report statistics: found via market_snapshots, found via positions, still NULL
- _Requirements: 2.3_
- [x] 7. Fix: Sentiment Z-Score Normalization
- [x] 7.1 Add normalize_impact_scores() function
- File: `services/aggregation/scoring.py` (new helper function)
- Query 7-day mean and stddev per ticker: `SELECT AVG(impact_score) as mean, STDDEV(impact_score) as stddev FROM document_impact_records WHERE ticker = $1 AND created_at >= NOW() - INTERVAL '7 days'`
- Formula: `normalized = (raw - mean_7d) / max(stddev_7d, 0.1)`
- Fallback: if fewer than 5 records in 7-day window, return raw score unchanged
- The 0.1 floor prevents division by near-zero stddev for low-activity tickers
- _Requirements: 2.4_
- [x] 7.2 Integrate normalization into aggregation loop
- File: `services/aggregation/worker.py`
- Before `compute_signal_weight()` call (around line 440), normalize `imp.impact_score` via `normalize_impact_scores()`
- Pass normalized value into `compute_signal_weight()` and `WeightedSignal`
- _Bug_Condition: input.impactScoreUsedRaw == TRUE AND input.ticker7dStddev > 0_
- _Expected_Behavior: Normalized impact scores with mean ≈ 0, stddev ≈ 1 for active tickers_
- _Preservation: Tickers with balanced distributions continue to produce neutral signals_
- _Requirements: 2.4, 3.3_
- [x] 8. Verify fixes pass all tests
- [x] 8.1 Verify bug condition exploration test now passes
- **Property 1: Expected Behavior** - Pipeline Health Bugs Resolved
- **IMPORTANT**: Re-run the SAME test from task 1 - do NOT write a new test
- The test from task 1 encodes the expected behavior for all five bugs
- Run `tests/test_pbt_pipeline_health_bug_condition.py`
- **EXPECTED OUTCOME**: Test PASSES (confirms bugs are fixed)
- _Requirements: 2.1, 2.2, 2.4, 2.6_
- [x] 8.2 Verify preservation tests still pass
- **Property 2: Preservation** - Pipeline Behavior Unchanged for Non-Bug Inputs
- **IMPORTANT**: Re-run the SAME tests from task 2 - do NOT write new tests
- Run `tests/test_pbt_pipeline_health_preservation.py`
- **EXPECTED OUTCOME**: Tests PASS (confirms no regressions)
- Confirm all preservation properties still hold after fixes
- _Requirements: 3.1, 3.2, 3.3, 3.5, 3.6, 3.7_
- [x] 9. Lint and final validation
- Run `.venv/bin/ruff check services/` and fix any lint errors
- Run `.venv/bin/python -m pytest tests/ -x --tb=short -q` to confirm full test suite passes
- Verify Helm template renders correctly with 0 signal-engine replicas
- _Requirements: all_
- [x] 10. Checkpoint - Ensure all tests pass
- Ensure all tests pass, ask the user if questions arise.
## Task Dependency Graph
```json
{
"waves": [
{"tasks": ["1", "2"]},
{"tasks": ["3", "4"]},
{"tasks": ["5"]},
{"tasks": ["6"]},
{"tasks": ["7"]},
{"tasks": ["8"]},
{"tasks": ["9"]},
{"tasks": ["10"]}
]
}
```
Tasks 3, 4 are independent config changes (no code deps).
Task 5 is independent but task 6 builds on the price fallback concept.
Task 7 is the most complex (new function + integration).
Tasks 8-10 must run after all fixes are applied.
## Notes
- Bug 4 (signal engine) is validated by Helm template rendering, not a unit test
- The backfill script (6.2) is a one-time operation, not covered by recurring tests
- Preservation tests use Hypothesis with `@settings(max_examples=100)` per project conventions
- Test files follow `test_pbt_*` naming convention per project standards