fix: pipeline health — stuck docs, price fallback, sentiment normalization, signal-engine scale, quality gate
- Scheduler: lower stale threshold 240→30 min, batch limit 100→500, TTL 14400→3600 - Prediction snapshot: add 24h market_snapshots time-window fallback - Aggregation: add normalize_impact_scores() z-score normalization - Helm: signal-engine replicas → 0 (idle when dual pipeline disabled) - Quality gate: max_snapshot_age_hours 24→48 - Add backfill script for NULL price_at_prediction snapshots - Add PBT bug condition and preservation tests (14 tests)
This commit is contained in:
@@ -0,0 +1 @@
|
||||
{"specId": "f5d99301-94ef-4dc2-8ba4-ccefeee7ecba", "workflowType": "requirements-first", "specType": "bugfix"}
|
||||
@@ -0,0 +1,51 @@
|
||||
# Bugfix Requirements Document
|
||||
|
||||
## Introduction
|
||||
|
||||
Five operational bugs in the stonks-beta deployment degrade pipeline health: 1,809 documents stuck in `parsed` status due to recovery batch limits, 26.5% of prediction snapshots missing prices due to incomplete fallback chains, 64% sell bias from uncalibrated NuExtract3 sentiment outputs, idle signal-engine consuming resources while doing nothing, and a quality gate stuck in paper-only mode due to an overly strict staleness threshold interacting with the NULL price problem.
|
||||
|
||||
## Bug Analysis
|
||||
|
||||
### Current Behavior (Defect)
|
||||
|
||||
1.1 WHEN the `recover_stale_documents` task runs with 1,809+ documents stuck in `parsed` status THEN the system only processes 100 per cycle (every ~5 minutes), requiring 90+ cycles (~7.5 hours) to clear the backlog while new documents may continue accumulating
|
||||
|
||||
1.2 WHEN a prediction snapshot is created for a ticker without an open position AND without recent market_snapshots data THEN the system stores NULL in `price_at_prediction` because the fallback chain stops at the positions table (26.5% of snapshots affected — 33,324 of 125,590)
|
||||
|
||||
1.3 WHEN the outcome evaluator encounters a prediction snapshot with NULL `price_at_prediction` THEN the system skips the snapshot entirely, creating a validation blind spot where 26.5% of predictions are never evaluated
|
||||
|
||||
1.4 WHEN the aggregation pipeline processes NuExtract3 extraction outputs THEN the system passes raw `impact_score` and `sentiment` values directly into signal weighting without any distribution normalization, resulting in systematic negative bias producing 64% sell / 23% watch / 12% buy recommendations
|
||||
|
||||
1.5 WHEN the signal-engine pod starts with `dual_pipeline_enabled=False` THEN the system enters an infinite sleep loop consuming CPU (100m request / 500m limit) and memory (128Mi request / 256Mi limit) while producing zero signal evaluations
|
||||
|
||||
1.6 WHEN the quality gate checks `model_metric_snapshots` freshness with a 24-hour staleness threshold AND the validation cycle skips all predictions due to NULL prices (Bug 1.2/1.3) THEN the system permanently defaults to paper-only mode because no fresh metric snapshots are ever generated
|
||||
|
||||
### Expected Behavior (Correct)
|
||||
|
||||
2.1 WHEN the scheduler detects more than 100 documents stuck in `parsed` status older than the threshold THEN the system SHALL increase the batch limit for recovery processing (up to 500 per cycle) and provide a one-time management command to bulk-recover the existing backlog without waiting for periodic sweeps
|
||||
|
||||
2.2 WHEN a prediction snapshot is created and no price is available from market_snapshots (exact time) or positions table THEN the system SHALL query `market_snapshots` with a wider time window (last 24 hours of bar data for the ticker) as an additional fallback before accepting NULL
|
||||
|
||||
2.3 WHEN backfilling existing prediction snapshots with NULL `price_at_prediction` THEN the system SHALL use the extended fallback chain (market_snapshots within 24h of `generated_at`, then positions) to populate prices retroactively via a migration script
|
||||
|
||||
2.4 WHEN the aggregation pipeline computes signal weights from impact records THEN the system SHALL apply z-score normalization to `impact_score` values relative to the rolling 7-day distribution of impact records for the same ticker, preventing systematic model bias from dominating the directional signal
|
||||
|
||||
2.5 WHEN the signal-engine deployment is not ready for production use (`dual_pipeline_enabled=False`) THEN the system SHALL be scaled to 0 replicas in the Helm values files (beta, paper, live) to eliminate wasted CPU, memory, and any GPU time-slice allocations
|
||||
|
||||
2.6 WHEN the quality gate evaluates metric snapshot freshness during the bootstrapping period THEN the system SHALL use a 48-hour staleness threshold (instead of 24h) to tolerate gaps while the validation cycle ramps up after Bug 1.2/1.3 are fixed
|
||||
|
||||
### Unchanged Behavior (Regression Prevention)
|
||||
|
||||
3.1 WHEN documents enter `parsed` status and are processed within the normal threshold window (< 240 minutes) THEN the system SHALL CONTINUE TO leave them for the extraction queue consumer without interference from the recovery task
|
||||
|
||||
3.2 WHEN a prediction snapshot is created and market_snapshots contains a recent bar for the ticker THEN the system SHALL CONTINUE TO use the primary `market_snapshots` close price without invoking any fallback
|
||||
|
||||
3.3 WHEN the aggregation pipeline processes tickers with balanced sentiment distributions (equal bullish/bearish evidence) THEN the system SHALL CONTINUE TO produce neutral/mixed recommendations without artificial skew from the normalization step
|
||||
|
||||
3.4 WHEN the signal-engine is re-enabled in the future (dual_pipeline_enabled=True with replicas > 0) THEN the system SHALL CONTINUE TO function correctly with its existing queue-based architecture and configuration loading
|
||||
|
||||
3.5 WHEN the quality gate evaluates a metric snapshot that is less than 48 hours old and meets all threshold criteria THEN the system SHALL CONTINUE TO promote recommendations to live_eligible mode per existing threshold logic
|
||||
|
||||
3.6 WHEN the outcome evaluator processes prediction snapshots with valid (non-NULL) prices THEN the system SHALL CONTINUE TO evaluate them normally and produce prediction_outcomes records
|
||||
|
||||
3.7 WHEN the `retry_failed_extractions` task handles documents in `extraction_failed` status THEN the system SHALL CONTINUE TO process them on the existing cadence and logic without interference from the parsed-document recovery changes
|
||||
@@ -0,0 +1,320 @@
|
||||
# Pipeline Health Fixes — Bugfix Design
|
||||
|
||||
## Overview
|
||||
|
||||
Five operational bugs degrade stonks-beta pipeline health. This design formalizes the bug conditions, expected fixes, and validation strategy for each:
|
||||
|
||||
1. **Stuck Parsed Docs** — `recover_stale_documents()` batch limit of 100 is too low for 1,809 stuck documents; increase to 500 and lower the stale threshold to 30 minutes.
|
||||
2. **Extended Price Fallback** — Prediction snapshots missing prices (26.5%) because the fallback chain stops at `positions`; add a third fallback querying `market_snapshots` within 24h.
|
||||
3. **Sentiment Z-Score Normalization** — Raw NuExtract3 `impact_score` values produce 64% sell bias; normalize using 7-day rolling z-scores per ticker before signal weighting.
|
||||
4. **Signal Engine Scale Down** — Idle signal-engine pods consume resources; set replicas to 0 in all Helm values files.
|
||||
5. **Quality Gate Threshold** — 24h staleness threshold permanently locks quality gate to paper-only; relax to 48h.
|
||||
|
||||
## Glossary
|
||||
|
||||
- **Bug_Condition (C)**: The specific conditions under which each bug manifests
|
||||
- **Property (P)**: The desired correct behavior after the fix is applied
|
||||
- **Preservation**: Existing behavior that must remain unchanged after the fix
|
||||
- **`recover_stale_documents()`**: Function in `services/scheduler/app.py` that re-enqueues documents stuck in `parsed` status
|
||||
- **`STALE_PARSED_THRESHOLD_MINUTES`**: Constant (currently 240) controlling how long a document must be stuck before recovery
|
||||
- **`fetch_latest_close_price()`**: Function in `services/validation/prediction_snapshot.py` that queries `market_snapshots` for the most recent bar
|
||||
- **`compute_signal_weight()`**: Function in `services/aggregation/scoring.py` that computes combined signal weight from recency, credibility, novelty, confidence, and impact
|
||||
- **`QualityGateConfig.max_snapshot_age_hours`**: Threshold in `services/trading/model_quality_gate.py` controlling when the quality gate defaults to paper-only
|
||||
|
||||
## Bug Details
|
||||
|
||||
### Bug Condition
|
||||
|
||||
The pipeline health degradation manifests across five independent conditions:
|
||||
|
||||
**Formal Specification:**
|
||||
```
|
||||
FUNCTION isBugCondition(input)
|
||||
INPUT: input of type PipelineState
|
||||
OUTPUT: boolean
|
||||
|
||||
-- Bug 1: Parsed docs stuck beyond batch capacity
|
||||
RETURN (input.stuckParsedDocCount > 100
|
||||
AND input.recoveryBatchLimit == 100
|
||||
AND input.docStaleMinutes >= 240)
|
||||
-- Bug 2: Price fallback chain incomplete
|
||||
OR (input.tickerPrice IS NULL
|
||||
AND input.positionPrice IS NULL
|
||||
AND input.marketSnapshotWithin24h IS NOT NULL)
|
||||
-- Bug 3: Raw impact scores without normalization
|
||||
OR (input.impactScoreUsedRaw == TRUE
|
||||
AND input.ticker7dStddev > 0)
|
||||
-- Bug 4: Signal engine running idle
|
||||
OR (input.signalEngineReplicas > 0
|
||||
AND input.dualPipelineEnabled == FALSE)
|
||||
-- Bug 5: Quality gate threshold too strict
|
||||
OR (input.snapshotAgeHours > 24
|
||||
AND input.snapshotAgeHours <= 48
|
||||
AND input.maxSnapshotAgeConfig == 24)
|
||||
END FUNCTION
|
||||
```
|
||||
|
||||
### Examples
|
||||
|
||||
- **Bug 1**: 1,809 documents in `parsed` status older than 4 hours. At 100/cycle every 5 minutes, clearing takes 90+ cycles (~7.5h). With 500/batch, it takes 4 cycles (~20 min).
|
||||
- **Bug 2**: Ticker PLTR has no open position and `fetch_latest_close_price` returns NULL, but `market_snapshots` has a bar from 3 hours ago that could serve as price.
|
||||
- **Bug 3**: NuExtract3 outputs `impact_score` values clustered around -0.3 to -0.1 for a ticker. Without normalization, `weighted_sentiment_average()` systematically produces negative signals → 64% sell recommendations.
|
||||
- **Bug 4**: signal-engine pod starts, detects `dual_pipeline_enabled=False`, enters infinite sleep loop consuming 100m CPU request / 128Mi memory request.
|
||||
- **Bug 5**: Quality gate reads `model_metric_snapshots`, finds the most recent is 26h old (because validation skips NULL-price predictions), fails staleness check, forces paper-only mode permanently.
|
||||
|
||||
## Expected Behavior
|
||||
|
||||
### Preservation Requirements
|
||||
|
||||
**Unchanged Behaviors:**
|
||||
- Documents entering `parsed` status and processed within the normal threshold window (< 30 min after fix) are left alone for the extraction queue consumer
|
||||
- Primary price lookup via `fetch_latest_close_price()` (exact time match from `market_snapshots`) continues as the first-choice price source
|
||||
- Tickers with balanced sentiment distributions continue to produce neutral/mixed recommendations without artificial skew
|
||||
- Signal-engine's queue-based architecture and configuration loading remain functional when re-enabled with replicas > 0
|
||||
- Quality gate threshold logic for snapshots younger than 48h and meeting all criteria continues to promote to `live_eligible`
|
||||
- Outcome evaluator continues to process prediction snapshots with valid (non-NULL) prices normally
|
||||
- `retry_failed_extractions` task continues on its existing cadence without interference
|
||||
|
||||
**Scope:**
|
||||
All inputs that do NOT match the bug conditions above should be completely unaffected by these fixes. The fixes are additive (new fallback path, wider batch, normalization layer) or config-only (replica count, threshold constant).
|
||||
|
||||
## Hypothesized Root Cause
|
||||
|
||||
### Bug 1: Stuck Parsed Docs
|
||||
- **Batch limit too small**: The `LIMIT 100` in the SQL query caps recovery throughput at 100 docs per scheduler cycle (~5 min). When a Redis crash orphans thousands of documents, the recovery rate cannot keep up with the backlog.
|
||||
- **Threshold too conservative**: `STALE_PARSED_THRESHOLD_MINUTES = 240` (4 hours) means documents must be stuck for 4 hours before recovery kicks in. A 30-minute threshold would catch orphans much faster.
|
||||
|
||||
### Bug 2: Incomplete Fallback Chain
|
||||
- **Missing time-window query**: `fetch_latest_close_price()` only checks `market_snapshots` for an exact timestamp match. When market data ingestion is delayed or the prediction happens outside market hours, no exact match exists.
|
||||
- **Positions-only fallback**: The positions table fallback only works for tickers with an active position. 26.5% of snapshots are for tickers without positions.
|
||||
|
||||
### Bug 3: Uncalibrated Impact Scores
|
||||
- **No distribution normalization**: NuExtract3 model outputs are passed raw into `compute_signal_weight()` via `impact_score` parameter. The model has a systematic negative bias in its output distribution that is not corrected.
|
||||
- **Per-ticker variance ignored**: Different tickers receive different volume/types of news, producing different impact_score distributions. A global normalization would be insufficient.
|
||||
|
||||
### Bug 4: Idle Signal Engine
|
||||
- **Replicas set to 1 by default**: `values.yaml` defines `signalEngine.replicas: 1` regardless of whether the dual pipeline feature is enabled. The pod starts, detects the feature is off, and sleeps forever.
|
||||
|
||||
### Bug 5: Overly Strict Staleness
|
||||
- **24h threshold too tight during bootstrapping**: The `max_snapshot_age_hours = 24` default assumes the validation cycle runs frequently. When Bug 2/3 cause most predictions to be skipped, metric snapshots aren't generated, and the 24h window expires.
|
||||
|
||||
## Correctness Properties
|
||||
|
||||
Property 1: Bug Condition — Stuck Parsed Docs Recovery
|
||||
|
||||
_For any_ set of documents stuck in `parsed` status longer than 30 minutes, the fixed `recover_stale_documents()` function SHALL process up to 500 documents per cycle, reducing backlog clearance time by 5x compared to the previous 100-document limit.
|
||||
|
||||
**Validates: Requirements 2.1**
|
||||
|
||||
Property 2: Bug Condition — Extended Price Fallback
|
||||
|
||||
_For any_ prediction snapshot where `fetch_latest_close_price()` returns NULL and the `positions` table has no price, but `market_snapshots` contains a bar for the ticker within 24 hours of the prediction time, the fixed code SHALL use that bar's close price as `price_at_prediction`.
|
||||
|
||||
**Validates: Requirements 2.2, 2.3**
|
||||
|
||||
Property 3: Bug Condition — Sentiment Z-Score Normalization
|
||||
|
||||
_For any_ set of `document_impact_records` for a ticker, the fixed aggregation pipeline SHALL normalize `impact_score` values using the 7-day rolling mean and standard deviation for that ticker before passing them into signal weight computation, preventing systematic model bias.
|
||||
|
||||
**Validates: Requirements 2.4**
|
||||
|
||||
Property 4: Bug Condition — Signal Engine Scale Down
|
||||
|
||||
_For any_ Helm deployment where `dual_pipeline_enabled=False`, the fixed Helm values SHALL specify `signalEngine.replicas: 0`, preventing the pod from being scheduled and consuming resources.
|
||||
|
||||
**Validates: Requirements 2.5**
|
||||
|
||||
Property 5: Bug Condition — Quality Gate Threshold
|
||||
|
||||
_For any_ model metric snapshot that is between 24h and 48h old, the fixed quality gate SHALL NOT reject it as stale, allowing the system to remain in non-paper mode during the bootstrapping period.
|
||||
|
||||
**Validates: Requirements 2.6**
|
||||
|
||||
Property 6: Preservation — Normal Document Processing
|
||||
|
||||
_For any_ document that enters `parsed` status and is processed within 30 minutes, the fixed `recover_stale_documents()` function SHALL NOT interfere with normal extraction queue processing, preserving the existing pipeline flow.
|
||||
|
||||
**Validates: Requirements 3.1, 3.7**
|
||||
|
||||
Property 7: Preservation — Primary Price Path
|
||||
|
||||
_For any_ prediction snapshot where `fetch_latest_close_price()` returns a valid price, the fixed code SHALL use that price directly without invoking any fallback, preserving the primary price lookup behavior.
|
||||
|
||||
**Validates: Requirements 3.2**
|
||||
|
||||
Property 8: Preservation — Balanced Sentiment
|
||||
|
||||
_For any_ ticker with a balanced sentiment distribution (equal bullish/bearish evidence), the z-score normalization SHALL produce values centered around 0, preserving neutral/mixed recommendation output without artificial skew.
|
||||
|
||||
**Validates: Requirements 3.3**
|
||||
|
||||
Property 9: Preservation — Quality Gate Valid Snapshots
|
||||
|
||||
_For any_ model metric snapshot younger than 48h that meets all threshold criteria, the fixed quality gate SHALL continue to promote recommendations to `live_eligible` mode per existing logic.
|
||||
|
||||
**Validates: Requirements 3.5, 3.6**
|
||||
|
||||
## Fix Implementation
|
||||
|
||||
### Changes Required
|
||||
|
||||
**Bug 1: Stuck Parsed Docs Recovery**
|
||||
|
||||
**File**: `services/scheduler/app.py`
|
||||
|
||||
**Function**: `recover_stale_documents()`
|
||||
|
||||
**Specific Changes**:
|
||||
1. **Lower stale threshold**: Change `STALE_PARSED_THRESHOLD_MINUTES` from `240` to `30` — documents stuck longer than 30 minutes are likely orphaned
|
||||
2. **Increase batch limit**: Change `LIMIT 100` to `LIMIT 500` in the SQL query
|
||||
3. **Update enqueued TTL**: Change `_ENQUEUED_TTL` from `14400` (4h) to `3600` (1h) to match the new threshold
|
||||
|
||||
---
|
||||
|
||||
**Bug 2: Extended Price Fallback**
|
||||
|
||||
**File**: `services/validation/prediction_snapshot.py`
|
||||
|
||||
**Function**: `create_prediction_snapshot()`
|
||||
|
||||
**Specific Changes**:
|
||||
1. **Add market_snapshots time-window fallback**: After the positions fallback fails, query `market_snapshots` for the most recent bar within 24h of the current time for the ticker
|
||||
2. **SQL query**: `SELECT close FROM market_snapshots WHERE ticker = $1 AND timestamp >= NOW() - INTERVAL '24 hours' ORDER BY timestamp DESC LIMIT 1`
|
||||
3. **Log the fallback**: Add info-level logging when the extended fallback is used
|
||||
|
||||
**New File**: `scripts/backfill_snapshot_prices.py`
|
||||
|
||||
**Purpose**: One-time backfill script to populate `price_at_prediction` for existing NULL snapshots using the extended fallback chain.
|
||||
|
||||
**Approach**:
|
||||
1. Query all `prediction_snapshots` where `price_at_prediction IS NULL`
|
||||
2. For each, attempt: `market_snapshots` within 24h of `generated_at`, then `positions` table
|
||||
3. Update the row with the found price
|
||||
4. Report statistics (found via market_snapshots, found via positions, still NULL)
|
||||
|
||||
---
|
||||
|
||||
**Bug 3: Sentiment Z-Score Normalization**
|
||||
|
||||
**File**: `services/aggregation/worker.py` (or new helper in `services/aggregation/scoring.py`)
|
||||
|
||||
**Function**: New function `normalize_impact_scores()` called before `compute_signal_weight()`
|
||||
|
||||
**Specific Changes**:
|
||||
1. **Add normalization function**: Compute 7-day rolling mean and stddev of `impact_score` per ticker from `document_impact_records`
|
||||
2. **Formula**: `normalized = (raw - mean_7d) / max(stddev_7d, 0.1)` — the 0.1 floor prevents division by near-zero stddev for low-activity tickers
|
||||
3. **Integration point**: In the aggregation loop (around line 440 of worker.py), normalize `imp.impact_score` before passing to `compute_signal_weight()` and `WeightedSignal`
|
||||
4. **Fallback**: If fewer than 5 records exist in the 7-day window, use the raw score (insufficient data for meaningful normalization)
|
||||
5. **Query**: `SELECT AVG(impact_score) as mean, STDDEV(impact_score) as stddev FROM document_impact_records WHERE ticker = $1 AND created_at >= NOW() - INTERVAL '7 days'`
|
||||
|
||||
---
|
||||
|
||||
**Bug 4: Signal Engine Scale Down**
|
||||
|
||||
**Files**: `infra/helm/stonks-oracle/values.yaml`, `values-beta.yaml`, `values-paper.yaml`
|
||||
|
||||
**Specific Changes**:
|
||||
1. **values.yaml**: Change `signalEngine.replicas` from `1` to `0`
|
||||
2. **values-beta.yaml**: Add `signalEngine.replicas: 0` under `services:`
|
||||
3. **values-paper.yaml**: Add `signalEngine.replicas: 0` under `services:`
|
||||
|
||||
---
|
||||
|
||||
**Bug 5: Quality Gate Threshold**
|
||||
|
||||
**File**: `services/trading/model_quality_gate.py`
|
||||
|
||||
**Class**: `QualityGateConfig`
|
||||
|
||||
**Specific Changes**:
|
||||
1. **Change default**: `max_snapshot_age_hours: int = 48` (was 24)
|
||||
|
||||
## Testing Strategy
|
||||
|
||||
### Validation Approach
|
||||
|
||||
The testing strategy follows a two-phase approach: first, surface counterexamples that demonstrate the bug on unfixed code, then verify the fix works correctly and preserves existing behavior.
|
||||
|
||||
### Exploratory Bug Condition Checking
|
||||
|
||||
**Goal**: Surface counterexamples that demonstrate the bugs BEFORE implementing the fixes. Confirm or refute the root cause analysis.
|
||||
|
||||
**Test Plan**: Write tests that exercise each bug condition on the unfixed code to observe failures.
|
||||
|
||||
**Test Cases**:
|
||||
1. **Batch Overflow Test**: Create 600 documents in `parsed` status older than threshold, run `recover_stale_documents()`, assert only 100 are processed (will demonstrate Bug 1)
|
||||
2. **Price Fallback Gap Test**: Call `create_prediction_snapshot()` for a ticker with no position and no exact market_snapshots match, assert `price_at_prediction` is NULL (will demonstrate Bug 2)
|
||||
3. **Sentiment Bias Test**: Generate 50 impact records with systematic negative bias (mean=-0.3, stddev=0.1), compute weighted signals, assert directional signal is negative (will demonstrate Bug 3)
|
||||
4. **Quality Gate Staleness Test**: Set most recent metric snapshot to 26h ago, evaluate quality gate, assert it fails (will demonstrate Bug 5)
|
||||
|
||||
**Expected Counterexamples**:
|
||||
- Bug 1: Only 100 of 600 documents recovered per cycle
|
||||
- Bug 2: `price_at_prediction` stored as NULL despite market data existing within 24h
|
||||
- Bug 3: Weighted sentiment average heavily negative despite mixed underlying events
|
||||
- Bug 5: Quality gate returns `passed=False` with reason containing "stale"
|
||||
|
||||
### Fix Checking
|
||||
|
||||
**Goal**: Verify that for all inputs where the bug condition holds, the fixed function produces the expected behavior.
|
||||
|
||||
**Pseudocode:**
|
||||
```
|
||||
FOR ALL input WHERE isBugCondition(input) DO
|
||||
result := fixedFunction(input)
|
||||
ASSERT expectedBehavior(result)
|
||||
END FOR
|
||||
```
|
||||
|
||||
**Per-bug fix checks:**
|
||||
- Bug 1: `recover_stale_documents()` processes up to 500 docs with 30-min threshold
|
||||
- Bug 2: Extended fallback returns a price when `market_snapshots` has data within 24h
|
||||
- Bug 3: Normalized impact scores have mean ≈ 0 and stddev ≈ 1 for active tickers
|
||||
- Bug 4: `kubectl get pods` shows 0 signal-engine pods
|
||||
- Bug 5: Quality gate passes for snapshots 24–48h old that meet metric thresholds
|
||||
|
||||
### Preservation Checking
|
||||
|
||||
**Goal**: Verify that for all inputs where the bug condition does NOT hold, the fixed function produces the same result as the original function.
|
||||
|
||||
**Pseudocode:**
|
||||
```
|
||||
FOR ALL input WHERE NOT isBugCondition(input) DO
|
||||
ASSERT originalFunction(input) = fixedFunction(input)
|
||||
END FOR
|
||||
```
|
||||
|
||||
**Testing Approach**: Property-based testing is recommended for preservation checking because:
|
||||
- It generates many test cases automatically across the input domain
|
||||
- It catches edge cases that manual unit tests might miss
|
||||
- It provides strong guarantees that behavior is unchanged for all non-buggy inputs
|
||||
|
||||
**Test Plan**: Observe behavior on UNFIXED code first for normal inputs, then write property-based tests capturing that behavior.
|
||||
|
||||
**Test Cases**:
|
||||
1. **Normal Doc Processing Preservation**: Documents < 30 min old are never touched by recovery
|
||||
2. **Primary Price Preservation**: When `fetch_latest_close_price()` succeeds, no fallback is invoked
|
||||
3. **Balanced Sentiment Preservation**: Tickers with symmetric impact_score distributions produce neutral signals after normalization
|
||||
4. **Quality Gate Normal Preservation**: Snapshots < 48h old and meeting thresholds still pass
|
||||
5. **Failed Extraction Preservation**: `retry_failed_extractions()` behavior unchanged
|
||||
|
||||
### Unit Tests
|
||||
|
||||
- Test `recover_stale_documents()` with various document counts (0, 50, 500, 1000)
|
||||
- Test extended price fallback with market_snapshots at various time offsets (1h, 12h, 23h, 25h)
|
||||
- Test z-score normalization with known distributions (mean=0, mean=-0.5, stddev=0, stddev=0.05)
|
||||
- Test quality gate with snapshot ages at boundary (23h, 24h, 47h, 48h, 49h)
|
||||
- Test backfill script with mixed NULL/non-NULL snapshots
|
||||
|
||||
### Property-Based Tests
|
||||
|
||||
- Generate random document ages and counts, verify recovery processes correct subset (> 30 min old, up to 500)
|
||||
- Generate random ticker price scenarios, verify fallback chain ordering is preserved (primary → positions → market_snapshots_24h → NULL)
|
||||
- Generate random impact_score distributions per ticker, verify normalized output has bounded variance and zero-centered mean
|
||||
- Generate random snapshot ages, verify quality gate accepts [0, 48h) and rejects [48h, ∞)
|
||||
|
||||
### Integration Tests
|
||||
|
||||
- End-to-end: create documents in `parsed` status, run scheduler cycle, verify extraction queue populated
|
||||
- End-to-end: create prediction snapshot for ticker without position, verify price populated from market_snapshots
|
||||
- End-to-end: run full aggregation cycle with biased NuExtract3 outputs, verify recommendation direction is not systematically biased
|
||||
- Helm template render: verify signal-engine deployment has 0 replicas in all value files
|
||||
@@ -0,0 +1,174 @@
|
||||
# Implementation Plan
|
||||
|
||||
## Overview
|
||||
|
||||
Bugfix implementation for five pipeline health issues: stuck parsed docs, missing price fallback, uncalibrated sentiment scores, idle signal-engine pods, and overly strict quality gate threshold. Tasks follow the exploratory bugfix workflow: explore bugs via tests, preserve existing behavior, implement fixes, validate.
|
||||
|
||||
## Tasks
|
||||
|
||||
- [x] 1. Write bug condition exploration test
|
||||
- **Property 1: Bug Condition** - Pipeline Health Degradation
|
||||
- **CRITICAL**: This test MUST FAIL on unfixed code - failure confirms the bugs exist
|
||||
- **DO NOT attempt to fix the test or the code when it fails**
|
||||
- **NOTE**: This test encodes the expected behavior - it will validate the fix when it passes after implementation
|
||||
- **GOAL**: Surface counterexamples that demonstrate all five bugs exist
|
||||
- **Scoped PBT Approach**: Scope properties to the concrete failing cases for each bug condition
|
||||
- Test file: `tests/test_pbt_pipeline_health_bug_condition.py`
|
||||
- **Bug 1 - Batch Overflow**: Create 600 documents in `parsed` status older than 30 min, run `recover_stale_documents()`, assert up to 500 are recovered per cycle (will FAIL on unfixed code which caps at 100)
|
||||
- **Bug 2 - Price Fallback Gap**: Call `create_prediction_snapshot()` for ticker with no position and no exact market_snapshots match but data within 24h exists, assert `price_at_prediction` is NOT NULL (will FAIL on unfixed code which returns NULL)
|
||||
- **Bug 3 - Sentiment Bias**: Generate 50 impact records with systematic negative bias (mean=-0.3, stddev=0.1), compute weighted signals via aggregation, assert normalized output is zero-centered (will FAIL on unfixed code which passes raw scores)
|
||||
- **Bug 5 - Quality Gate Staleness**: Set most recent metric snapshot to 26h ago, evaluate quality gate, assert it passes (will FAIL on unfixed code which rejects at 24h)
|
||||
- Run tests on UNFIXED code
|
||||
- **EXPECTED OUTCOME**: Tests FAIL (this is correct - it proves the bugs exist)
|
||||
- Document counterexamples: batch capped at 100, price stored as NULL, sentiment heavily negative, quality gate returns `passed=False`
|
||||
- Mark task complete when tests are written, run, and failures are documented
|
||||
- _Requirements: 1.1, 1.2, 1.3, 1.4, 1.6_
|
||||
|
||||
- [x] 2. Write preservation property tests (BEFORE implementing fix)
|
||||
- **Property 2: Preservation** - Pipeline Behavior Unchanged for Non-Bug Inputs
|
||||
- **IMPORTANT**: Follow observation-first methodology
|
||||
- Test file: `tests/test_pbt_pipeline_health_preservation.py`
|
||||
- **Normal Doc Processing**: Observe that documents < 30 min old are never touched by `recover_stale_documents()` on unfixed code. Write property: for all documents with age < 30 min, recovery task does NOT enqueue them.
|
||||
- **Primary Price Path**: Observe that when `fetch_latest_close_price()` returns a valid price, no fallback is invoked. Write property: for all tickers where primary price exists, result equals primary price.
|
||||
- **Balanced Sentiment**: Observe that tickers with symmetric impact_score distributions (mean ≈ 0) produce neutral signals. Write property: for all impact_score sets with mean ≈ 0, normalized output remains centered around 0.
|
||||
- **Quality Gate Normal**: Observe that snapshots < 48h old meeting thresholds pass the quality gate. Write property: for all snapshot ages in [0, 48h) meeting metric criteria, quality gate returns `passed=True`.
|
||||
- **Failed Extraction Independence**: Observe `retry_failed_extractions()` behavior is unaffected. Write property: for all documents in `extraction_failed` status, retry logic unchanged.
|
||||
- Run tests on UNFIXED code
|
||||
- **EXPECTED OUTCOME**: Tests PASS (this confirms baseline behavior to preserve)
|
||||
- Mark task complete when tests are written, run, and passing on unfixed code
|
||||
- _Requirements: 3.1, 3.2, 3.3, 3.5, 3.6, 3.7_
|
||||
|
||||
- [x] 3. Fix: Signal Engine Scale Down (Helm values)
|
||||
|
||||
- [x] 3.1 Set signal-engine replicas to 0 in all Helm values files
|
||||
- In `infra/helm/stonks-oracle/values.yaml`: change `signalEngine.replicas` from `1` to `0`
|
||||
- In `infra/helm/stonks-oracle/values-beta.yaml`: add/set `signalEngine.replicas: 0` under `services:`
|
||||
- In `infra/helm/stonks-oracle/values-paper.yaml`: add/set `signalEngine.replicas: 0` under `services:`
|
||||
- _Bug_Condition: input.signalEngineReplicas > 0 AND input.dualPipelineEnabled == FALSE_
|
||||
- _Expected_Behavior: signalEngine.replicas == 0 when dual pipeline disabled_
|
||||
- _Preservation: Signal-engine architecture remains functional when re-enabled with replicas > 0_
|
||||
- _Requirements: 2.5, 3.4_
|
||||
|
||||
- [x] 4. Fix: Quality Gate Threshold Relaxation
|
||||
|
||||
- [x] 4.1 Change max_snapshot_age_hours default from 24 to 48
|
||||
- File: `services/trading/model_quality_gate.py`
|
||||
- In `QualityGateConfig` class, change `max_snapshot_age_hours: int = 24` to `max_snapshot_age_hours: int = 48`
|
||||
- _Bug_Condition: input.snapshotAgeHours > 24 AND input.snapshotAgeHours <= 48 AND input.maxSnapshotAgeConfig == 24_
|
||||
- _Expected_Behavior: Quality gate accepts snapshots up to 48h old_
|
||||
- _Preservation: Snapshots < 48h meeting criteria continue to promote to live_eligible_
|
||||
- _Requirements: 2.6, 3.5_
|
||||
|
||||
- [x] 5. Fix: Stuck Parsed Docs Recovery
|
||||
|
||||
- [x] 5.1 Lower STALE_PARSED_THRESHOLD_MINUTES from 240 to 30
|
||||
- File: `services/scheduler/app.py`
|
||||
- Change constant: `STALE_PARSED_THRESHOLD_MINUTES = 30`
|
||||
- Documents stuck longer than 30 minutes are likely orphaned
|
||||
- _Requirements: 2.1_
|
||||
|
||||
- [x] 5.2 Increase recovery batch LIMIT from 100 to 500
|
||||
- File: `services/scheduler/app.py`
|
||||
- In `recover_stale_documents()` SQL query, change `LIMIT 100` to `LIMIT 500`
|
||||
- _Requirements: 2.1_
|
||||
|
||||
- [x] 5.3 Update _ENQUEUED_TTL from 14400 to 3600
|
||||
- File: `services/scheduler/app.py`
|
||||
- Change `_ENQUEUED_TTL = 3600` (1 hour, matching the new recovery cadence)
|
||||
- _Bug_Condition: input.stuckParsedDocCount > 100 AND input.recoveryBatchLimit == 100 AND input.docStaleMinutes >= 240_
|
||||
- _Expected_Behavior: Recovery processes up to 500 docs per cycle with 30-min threshold_
|
||||
- _Preservation: Documents < 30 min old left alone for extraction queue consumer_
|
||||
- _Requirements: 2.1, 3.1, 3.7_
|
||||
|
||||
- [x] 6. Fix: Extended Price Fallback
|
||||
|
||||
- [x] 6.1 Add market_snapshots 24h time-window fallback to create_prediction_snapshot()
|
||||
- File: `services/validation/prediction_snapshot.py`
|
||||
- After positions fallback fails, query: `SELECT close FROM market_snapshots WHERE ticker = $1 AND timestamp >= NOW() - INTERVAL '24 hours' ORDER BY timestamp DESC LIMIT 1`
|
||||
- Add info-level logging when extended fallback is used
|
||||
- _Bug_Condition: input.tickerPrice IS NULL AND input.positionPrice IS NULL AND input.marketSnapshotWithin24h IS NOT NULL_
|
||||
- _Expected_Behavior: Use market_snapshots bar close price as price_at_prediction_
|
||||
- _Preservation: Primary fetch_latest_close_price() path unchanged when it returns a valid price_
|
||||
- _Requirements: 2.2, 3.2_
|
||||
|
||||
- [x] 6.2 Create backfill script scripts/backfill_snapshot_prices.py
|
||||
- Query all `prediction_snapshots` where `price_at_prediction IS NULL`
|
||||
- For each, attempt: `market_snapshots` within 24h of `generated_at`, then `positions` table
|
||||
- Update row with found price
|
||||
- Report statistics: found via market_snapshots, found via positions, still NULL
|
||||
- _Requirements: 2.3_
|
||||
|
||||
- [x] 7. Fix: Sentiment Z-Score Normalization
|
||||
|
||||
- [x] 7.1 Add normalize_impact_scores() function
|
||||
- File: `services/aggregation/scoring.py` (new helper function)
|
||||
- Query 7-day mean and stddev per ticker: `SELECT AVG(impact_score) as mean, STDDEV(impact_score) as stddev FROM document_impact_records WHERE ticker = $1 AND created_at >= NOW() - INTERVAL '7 days'`
|
||||
- Formula: `normalized = (raw - mean_7d) / max(stddev_7d, 0.1)`
|
||||
- Fallback: if fewer than 5 records in 7-day window, return raw score unchanged
|
||||
- The 0.1 floor prevents division by near-zero stddev for low-activity tickers
|
||||
- _Requirements: 2.4_
|
||||
|
||||
- [x] 7.2 Integrate normalization into aggregation loop
|
||||
- File: `services/aggregation/worker.py`
|
||||
- Before `compute_signal_weight()` call (around line 440), normalize `imp.impact_score` via `normalize_impact_scores()`
|
||||
- Pass normalized value into `compute_signal_weight()` and `WeightedSignal`
|
||||
- _Bug_Condition: input.impactScoreUsedRaw == TRUE AND input.ticker7dStddev > 0_
|
||||
- _Expected_Behavior: Normalized impact scores with mean ≈ 0, stddev ≈ 1 for active tickers_
|
||||
- _Preservation: Tickers with balanced distributions continue to produce neutral signals_
|
||||
- _Requirements: 2.4, 3.3_
|
||||
|
||||
- [x] 8. Verify fixes pass all tests
|
||||
|
||||
- [x] 8.1 Verify bug condition exploration test now passes
|
||||
- **Property 1: Expected Behavior** - Pipeline Health Bugs Resolved
|
||||
- **IMPORTANT**: Re-run the SAME test from task 1 - do NOT write a new test
|
||||
- The test from task 1 encodes the expected behavior for all five bugs
|
||||
- Run `tests/test_pbt_pipeline_health_bug_condition.py`
|
||||
- **EXPECTED OUTCOME**: Test PASSES (confirms bugs are fixed)
|
||||
- _Requirements: 2.1, 2.2, 2.4, 2.6_
|
||||
|
||||
- [x] 8.2 Verify preservation tests still pass
|
||||
- **Property 2: Preservation** - Pipeline Behavior Unchanged for Non-Bug Inputs
|
||||
- **IMPORTANT**: Re-run the SAME tests from task 2 - do NOT write new tests
|
||||
- Run `tests/test_pbt_pipeline_health_preservation.py`
|
||||
- **EXPECTED OUTCOME**: Tests PASS (confirms no regressions)
|
||||
- Confirm all preservation properties still hold after fixes
|
||||
- _Requirements: 3.1, 3.2, 3.3, 3.5, 3.6, 3.7_
|
||||
|
||||
- [x] 9. Lint and final validation
|
||||
- Run `.venv/bin/ruff check services/` and fix any lint errors
|
||||
- Run `.venv/bin/python -m pytest tests/ -x --tb=short -q` to confirm full test suite passes
|
||||
- Verify Helm template renders correctly with 0 signal-engine replicas
|
||||
- _Requirements: all_
|
||||
|
||||
- [x] 10. Checkpoint - Ensure all tests pass
|
||||
- Ensure all tests pass, ask the user if questions arise.
|
||||
|
||||
## Task Dependency Graph
|
||||
|
||||
```json
|
||||
{
|
||||
"waves": [
|
||||
{"tasks": ["1", "2"]},
|
||||
{"tasks": ["3", "4"]},
|
||||
{"tasks": ["5"]},
|
||||
{"tasks": ["6"]},
|
||||
{"tasks": ["7"]},
|
||||
{"tasks": ["8"]},
|
||||
{"tasks": ["9"]},
|
||||
{"tasks": ["10"]}
|
||||
]
|
||||
}
|
||||
```
|
||||
|
||||
Tasks 3, 4 are independent config changes (no code deps).
|
||||
Task 5 is independent but task 6 builds on the price fallback concept.
|
||||
Task 7 is the most complex (new function + integration).
|
||||
Tasks 8-10 must run after all fixes are applied.
|
||||
|
||||
## Notes
|
||||
|
||||
- Bug 4 (signal engine) is validated by Helm template rendering, not a unit test
|
||||
- The backfill script (6.2) is a one-time operation, not covered by recurring tests
|
||||
- Preservation tests use Hypothesis with `@settings(max_examples=100)` per project conventions
|
||||
- Test files follow `test_pbt_*` naming convention per project standards
|
||||
Reference in New Issue
Block a user