30 KiB
Design Document: Math Core v3 Engine
Overview
The v3 Calibrated Evidence Engine replaces the current dual-mode pipeline (heuristic + probabilistic) with a principled Bayesian evidence accumulation system. The upgrade transforms the signal processing core from weighted-sentiment averaging to:
EvidenceUnit → calibrated reliability (q_i) → log-likelihood ratio (LLR_i)
→ correlation-adjusted cluster LLR → posterior P_up → return distribution → EV/risk decision
Key design goals:
- Replace arbitrary weight products with calibrated probabilistic evidence
- Prevent correlated articles from inflating evidence counts
- Make confidence multiplicative so one bad dimension suppresses the trade
- Size positions using fractional Kelly criterion under hard risk caps
- Preserve the existing three-layer architecture and service boundaries
- Gate behind
v3_engine_enabledfeature flag with heuristic fallback
What changes vs. what stays:
- The
WeightedSignalabstraction remains as an intermediate before LLR conversion - All output goes into existing JSONB metadata columns (no new migrations)
- Service file boundaries are preserved; internals are upgraded
- The heuristic pipeline stays as fallback, controlled by a single flag read per aggregation cycle
Architecture
flowchart TD
subgraph Input Layer
CS[Company Signals]
MS[Macro Signals]
XS[Competitive Signals]
end
subgraph Normalization
EU[EvidenceUnit Normalization]
end
subgraph Scoring Pipeline
QI[Calibrated Reliability q_i]
LLR[LLR Conversion]
end
subgraph Clustering
CL[Correlation-Aware Clustering]
NEFF[n_eff Computation]
CLLR[Cluster LLR]
end
subgraph Posterior Assembly
REG[Regime Detection v3]
POST[Log-Odds Posterior P_up]
CON[LLR Entropy Contradiction]
CONF[Multiplicative Confidence]
end
subgraph Decision Layer
PROJ[Posterior State Projection]
RET[Return Distribution / EV Gate]
ELIG[Regime-Aware Eligibility]
KELLY[Fractional Kelly Sizing]
STOP[Regime-Aware Stops]
HEAT[Stop-Defined Portfolio Heat]
end
subgraph Quality & Control
DQ[Data Quality v3]
FF[Feature Flag Router]
TIER[Risk Tier Auto-Adjustment]
end
CS --> EU
MS --> EU
XS --> EU
EU --> QI
QI --> LLR
LLR --> CL
CL --> NEFF
NEFF --> CLLR
REG --> POST
CLLR --> POST
POST --> CON
POST --> CONF
CONF --> PROJ
PROJ --> RET
RET --> ELIG
ELIG --> KELLY
KELLY --> STOP
KELLY --> HEAT
DQ --> CONF
FF --> EU
TIER --> KELLY
Service File Mapping
| Service File | v3 Responsibility |
|---|---|
services/aggregation/scoring.py |
EvidenceUnit normalization, q_i pipeline, LLR conversion |
services/aggregation/bayesian.py |
Posterior assembly via log-odds, P_up, strength |
services/aggregation/contradiction.py |
LLR entropy contradiction score |
services/aggregation/regime.py |
Regime detection v3 (ATR-normalized trend_z) |
services/aggregation/interpolation.py |
Noisy-OR macro exposure, LLR emission |
services/aggregation/signal_propagation.py |
Correlation-shrunk competitive propagation |
services/aggregation/projection.py |
Posterior state A_t, regime-aware decay |
services/aggregation/worker.py |
Orchestration, clustering, n_eff, confidence assembly |
services/recommendation/eligibility.py |
EV gate, regime-aware eligibility, mode escalation |
services/trading/position_sizer.py |
Fractional Kelly sizing under caps |
services/trading/stop_loss_manager.py |
Regime-aware stop/TP, trailing activation |
services/risk/engine.py |
Stop-defined heat, tier auto-adjustment |
Feature Flag Flow
sequenceDiagram
participant Worker as aggregation/worker.py
participant DB as risk_configs table
participant V3 as v3 Pipeline
participant Heuristic as Heuristic Pipeline
Worker->>DB: SELECT v3_engine_enabled
alt v3_engine_enabled = True
Worker->>V3: Run v3 pipeline
V3-->>Worker: Posterior + confidence + EV
alt Unhandled error
V3-->>Worker: Exception
Worker->>Heuristic: Fallback to heuristic
Worker->>Worker: Log error, record fallback in metadata
end
else v3_engine_enabled = False or DB error
Worker->>Heuristic: Run heuristic pipeline
end
Components and Interfaces
1. EvidenceUnit Dataclass (scoring.py)
@dataclass(frozen=True)
class EvidenceUnit:
"""Canonical normalized signal representation for v3 pipeline."""
symbol: str
layer: str # "company" | "macro" | "competitive"
event_type: str
source_id: str
source_group: str
timestamp: datetime
horizon: str # "intraday" | "1d" | "7d" | "30d" | "90d"
direction: int # -1, 0, +1
sentiment_strength: float # [0, 1]
impact: float # [0, 1]
extraction_conf: float # [0, 1]
source_cred: float # [0, 1]
novelty: float # [0, 1]
event_base_rate: float # (0, 1]
cluster_id: str
Normalization functions — one per layer:
normalize_company_signal(impact_row, ...) -> EvidenceUnitnormalize_macro_signal(macro_impact_record, global_event, ...) -> EvidenceUnitnormalize_competitive_signal(competitive_signal_record, ...) -> EvidenceUnit
Each validates required fields (symbol, timestamp, source_id), substitutes 0.5 for missing optional numeric fields, and assigns direction from sentiment/impact_direction strings.
2. Calibrated Reliability Pipeline (scoring.py)
@dataclass(frozen=True)
class ReliabilityComponents:
q_ext: float
q_source: float
q_recency: float
q_uniqueness: float
q_i: float # final combined reliability
def compute_v3_reliability(
unit: EvidenceUnit,
source_stats: SourceStats,
cluster_position: int, # duplicate_count_before
reference_time: datetime,
) -> ReliabilityComponents: ...
Sub-computations:
q_ext = sigmoid(8.0 * (extraction_conf - 0.55))q_source = clamp((E[theta_s] - 0.50) / 0.35, 0, 1)with Beta(alpha_0+hits, beta_0+misses)q_recency = 2^(-age_hours / tau_adaptive)with adaptive half-lifeq_uniqueness = clamp(0.5 + 0.5 * novelty, 0.5, 1.0) * (1 / sqrt(1 + dup_count))q_i = clamp(q_ext * q_source * source_cred * q_recency * q_uniqueness, 0, 1)
3. LLR Conversion (scoring.py)
def compute_llr(unit: EvidenceUnit, q_i: float) -> float:
"""Convert calibrated reliability to log-likelihood ratio."""
p_correct = clamp(0.50 + 0.35 * q_i * unit.impact * unit.sentiment_strength, 0.501, 0.85)
if unit.direction == 0:
return 0.0
return unit.direction * math.log(p_correct / (1 - p_correct))
4. Correlation-Aware Clustering (worker.py)
@dataclass
class EvidenceCluster:
cluster_id: str
units: list[EvidenceUnit]
llrs: list[float]
n_eff: float
cluster_llr: float
def cluster_evidence(units: list[EvidenceUnit], llrs: list[float]) -> list[EvidenceCluster]:
"""Group by (symbol, horizon, event_type, source_group, time_bucket)."""
...
def compute_n_eff(llrs: list[float], correlations: list[list[float]]) -> float:
"""n_eff = (sum w_i)^2 / (sum w_i^2 + 2*sum_{i<j} rho_ij*w_i*w_j)"""
...
def compute_cluster_llr(llrs: list[float], n_eff: float) -> float:
"""LLR_c = clamp(weighted_mean(LLR_i, |LLR_i|) * sqrt(n_eff), -2.5, 2.5)"""
...
Default pairwise correlations:
| Relationship | rho |
|---|---|
| Same wire/story/source group | 0.80 |
| Same event, different publisher | 0.50 |
| Same theme, different event | 0.25 |
| Independent events | 0.00 |
5. Posterior Assembly (bayesian.py)
@dataclass(frozen=True)
class V3Posterior:
p_up: float # sigmoid(log_odds)
p_down: float # 1 - p_up
log_odds: float # logit(P_prior) + sum(gamma * LLR_c)
strength: float # abs(2 * p_up - 1)
direction: str # bullish | bearish | neutral
n_eff_total: float
regime: str
def compute_v3_posterior(
clusters: list[EvidenceCluster],
regime: RegimeClassification,
p_prior: float = 0.50,
) -> V3Posterior: ...
Direction thresholds by regime:
| Regime | Bullish if P_up >= | Bearish if P_up <= |
|---|---|---|
| panic | 0.68 | 0.32 |
| trend_following | 0.60 | 0.40 |
| mean_reversion | 0.63 | 0.37 |
| uncertainty | 0.65 | 0.35 |
6. Regime Detection v3 (regime.py)
@dataclass(frozen=True)
class V3RegimeClassification:
regime: MarketRegime
trend_z: float # (EMA_20 - EMA_100) / ATR_20
vol_ratio: float # sigma_20 / sigma_100
evidence_multiplier: float # gamma_regime
confidence_multiplier: float
phi_decay: float # for projection
atr_multiplier: float # for stops
def classify_regime_v3(
closing_prices: list[float],
daily_returns: list[float],
atr_20: float,
) -> V3RegimeClassification: ...
Classification rules (in priority order):
- Panic: vol_ratio > 1.5 OR abs(trend_z) > 2.5
- Trend following: abs(trend_z) >= 0.75 AND vol_ratio < 1.3
- Mean reversion: abs(trend_z) < 0.50 AND vol_ratio < 1.0
- Uncertainty: all other cases
7. LLR Entropy Contradiction (contradiction.py)
def compute_v3_contradiction(clusters: list[EvidenceCluster]) -> float:
"""
E_pos = sum(max(LLR_c, 0))
E_neg = sum(max(-LLR_c, 0))
H = -f_pos*log2(f_pos) - f_neg*log2(f_neg)
volume_factor = 1 - exp(-E_total / 3.0)
return H * volume_factor
"""
...
8. Multiplicative Confidence (worker.py)
def compute_v3_confidence(
n_eff_total: float,
q_values: list[float],
llrs: list[float],
strength: float,
regime_confidence_mult: float,
contradiction: float,
data_quality: float,
) -> float:
"""
C_evidence = 1 - exp(-n_eff_total / 5.0)
C_quality = weighted_mean(q_i, |LLR_i|)
confidence = clamp(
C_evidence * sqrt(C_quality) * sqrt(max(strength, 0.05))
* regime_confidence_mult * (1 - contradiction) * data_quality,
0, 1
)
"""
...
9. Noisy-OR Macro Exposure (interpolation.py)
def compute_normalized_macro_exposure(overlaps: dict[str, float]) -> float:
"""
E_raw = 1 - prod(1 - w_k * O_k)
E_max = 1 - prod(1 - w_k)
return E_raw / E_max
"""
...
def compute_macro_llr(
macro_impact: float,
event_confidence: float,
q_recency: float,
macro_direction: int,
) -> float:
"""
p_macro = clamp(0.50 + 0.30 * macro_impact * event_confidence * q_recency, 0.501, 0.80)
return macro_direction * log(p_macro / (1 - p_macro))
"""
...
10. Correlation-Shrunk Competitive Propagation (signal_propagation.py)
def compute_shrunk_correlation(
rho_rolling: float,
n_observations: int,
same_sector: bool,
) -> float:
"""
rho_prior = 0.30 if same_sector else 0.10
rho_shrunk = (n/(n+30)) * rho_rolling + (30/(n+30)) * rho_prior
return max(rho_shrunk, 0)
"""
...
def compute_competitive_llr(
llr_source: float,
rho_effective: float,
d_network: int,
pattern_confidence: float,
) -> float:
"""
attenuation = rho_effective * exp(-0.85 * d_network)
return clamp(llr_source * attenuation * pattern_confidence, -1.25, 1.25)
"""
...
11. Posterior State Projection (projection.py)
@dataclass
class V3ProjectionState:
a_t: float # accumulated evidence state
p_up_projected: float # sigmoid(logit(P_prior) + phi^h * A_t)
projected_strength: float
diverges: bool
phi_regime: float
def compute_v3_projection(
a_prev: float,
cluster_llrs: list[float],
regime: V3RegimeClassification,
p_prior: float,
projection_horizon: int,
known_catalyst_llr: float = 0.0,
) -> V3ProjectionState: ...
Regime decay factors (phi):
| Regime | phi |
|---|---|
| panic | 0.35 |
| trend_following | 0.80 |
| mean_reversion | 0.55 |
| uncertainty | 0.50 |
12. Return Distribution and EV Gate (eligibility.py)
@dataclass(frozen=True)
class ReturnDistribution:
sigma_h: float # realized_vol_20d * sqrt(horizon_days / 252)
mu_h: float # tanh(A_projected / 3.0) * confidence * sigma_h
ev_long: float # mu_h - costs - 0.10 * CVaR_5
min_edge: float # regime-specific minimum edge
eligible: bool
def compute_return_distribution(
a_projected: float,
confidence: float,
realized_vol_20d: float,
horizon_days: int,
costs: float,
regime: str,
) -> ReturnDistribution: ...
13. Fractional Kelly Position Sizing (position_sizer.py)
def compute_kelly_sizing(
p_win: float, # P_up from posterior
b: float, # reward ratio: clamp(1.2 + 2*conf + str - contra, 1.2, 3.0)
confidence: float,
data_quality: float,
contradiction: float,
max_position_pct: float,
available_caps: dict[str, float], # sector, correlation, heat capacities
) -> float:
"""
f_kelly = (p_win * b - (1 - p_win)) / b
portfolio_pct = clamp(max(0, f_kelly) * 0.25 * confidence * data_quality * (1 - contradiction), 0, max_position_pct)
Apply min of all capacity constraints.
"""
...
14. Stop-Defined Portfolio Heat (risk/engine.py)
def compute_portfolio_heat(
positions: list[OpenPosition],
stop_distances: dict[str, float],
) -> float:
"""risk_dollars = position_value * stop_distance_pct; heat = sum(risk_dollars)"""
...
def check_heat_capacity(
current_heat: float,
new_risk_dollars: float,
max_heat_pct: float,
portfolio_value: float,
) -> bool: ...
15. Data Quality v3 (worker.py)
def compute_v3_data_quality(
units: list[EvidenceUnit],
extraction_failure_rate: float,
age_newest_hours: float,
n_source_types: int,
) -> float:
"""
Q_parse = 1 - extraction_failure_rate
Q_conf = weighted_mean(extraction_conf, impact)
Q_fresh = exp(-age_newest_hours / 168)
Q_coverage = 1 - exp(-N_valid / 5)
Q_diversity = min(1, log2(1 + N_source_types) / log2(4))
return clamp(Q_parse * sqrt(Q_conf) * Q_fresh * Q_coverage * Q_diversity, 0, 1)
"""
...
16. Risk Tier Auto-Adjustment (risk/engine.py)
@dataclass
class TierMetrics:
profit_factor_30d: float
max_drawdown_30d: float
calibration_error: float
realized_sharpe_30d: float
n_trades_30d: int
reserve_pool_pct: float
def evaluate_tier_adjustment(metrics: TierMetrics) -> str:
"""Returns 'upgrade' | 'downgrade' | 'hold'"""
...
Data Models
EvidenceUnit (frozen dataclass)
| Field | Type | Range | Source |
|---|---|---|---|
| symbol | str | — | Required from signal |
| layer | str | company/macro/competitive | Set during normalization |
| event_type | str | — | catalyst_type or impact_type |
| source_id | str | — | document_id or event_id |
| source_group | str | — | publisher / "macro" / "competitive" |
| timestamp | datetime | — | published_at |
| horizon | str | intraday/1d/7d/30d/90d | window or estimated_duration mapping |
| direction | int | -1, 0, +1 | sentiment/direction mapping |
| sentiment_strength | float | [0, 1] | impact_score or sentiment confidence |
| impact | float | [0, 1] | impact_score |
| extraction_conf | float | [0, 1] | confidence field |
| source_cred | float | [0, 1] | source_credibility |
| novelty | float | [0, 1] | novelty_score |
| event_base_rate | float | (0, 1] | lookup by event_type |
| cluster_id | str | — | computed hash of clustering key |
V3 Posterior Output (stored in JSONB metadata)
{
"v3_posterior": {
"p_up": 0.64,
"p_down": 0.36,
"log_odds": 0.58,
"strength": 0.28,
"confidence": 0.61,
"contradiction": 0.18,
"n_eff": 5.7,
"data_quality": 0.82,
"regime": "trend_following"
},
"v3_return_model": {
"mu_h": 0.012,
"sigma_h": 0.041,
"ev_long": 0.007,
"min_edge": 0.0035
},
"v3_explainability": {
"top_positive_clusters": [...],
"top_negative_clusters": [...],
"suppression_reasons": [],
"risk_adjustments": []
},
"pipeline_mode": "v3"
}
Source Statistics (for q_source computation)
@dataclass
class SourceStats:
source_id: str
hits: int = 0 # correct directional predictions
misses: int = 0 # incorrect directional predictions
alpha_0: int = 3 # prior
beta_0: int = 3 # prior
Regime Parameters Table
| Regime | gamma (evidence) | confidence_mult | phi (decay) | ATR_mult (stops) | min_edge |
|---|---|---|---|---|---|
| panic | 0.70 | 0.70 | 0.35 | 2.5 | 0.0100 |
| trend_following | 1.10 | 1.00 | 0.80 | 1.8 | 0.0035 |
| mean_reversion | 0.90 | 0.95 | 0.55 | 1.4 | 0.0050 |
| uncertainty | 0.80 | 0.85 | 0.50 | 2.0 | 0.0075 |
Eligibility Thresholds (Regime-Specific)
| Regime | confidence_min | contradiction_max | strength_min |
|---|---|---|---|
| panic | 0.70 | 0.25 | 0.36 |
| trend_following | 0.55 | 0.40 | 0.20 |
| mean_reversion | 0.60 | 0.35 | 0.26 |
| uncertainty | 0.65 | 0.30 | 0.30 |
Correctness Properties
A property is a characteristic or behavior that should hold true across all valid executions of a system — essentially, a formal statement about what the system should do. Properties serve as the bridge between human-readable specifications and machine-verifiable correctness guarantees.
Property 1: Reliability q_i is bounded in [0, 1]
For any valid EvidenceUnit with extraction_conf in [0,1], source_cred in [0,1], novelty in [0,1], any non-negative age_hours, and any non-negative duplicate_count_before, the computed q_i SHALL be in the range [0.0, 1.0].
Validates: Requirements 2.8, 21.1
Property 2: p_correct is bounded in [0.501, 0.85]
For any valid q_i in [0, 1], impact in [0, 1], and sentiment_strength in [0, 1], the computed p_correct SHALL be in the range [0.501, 0.85].
Validates: Requirements 3.1, 21.2
Property 3: LLR sign matches direction and magnitude is bounded
For any valid signal with direction in {-1, +1}, the computed LLR SHALL have the same sign as direction, with absolute magnitude in [ln(0.501/0.499), ln(0.85/0.15)] ≈ [0.004, 1.735].
Validates: Requirements 3.2, 3.4, 3.5, 3.6, 21.3
Property 4: Neutral signals produce zero LLR
For any valid EvidenceUnit with direction = 0, regardless of all other field values, the computed LLR SHALL be exactly 0.0.
Validates: Requirements 1.5, 3.3
Property 5: Effective evidence count n_eff is bounded by cluster size
For any cluster of N signals with non-negative pairwise correlations rho_ij in [0, 1], the computed n_eff SHALL satisfy 0 < n_eff <= N.
Validates: Requirements 4.2, 21.4
Property 6: Cluster LLR is clamped to [-2.5, 2.5]
For any cluster configuration with any number of signals and any LLR values, the computed cluster LLR_c SHALL be in the range [-2.5, 2.5].
Validates: Requirements 4.4, 4.5
Property 7: Posterior P_up is in open interval (0, 1)
For any set of cluster LLRs (each in [-2.5, 2.5]), any regime evidence multiplier gamma in {0.70, 0.80, 0.90, 1.10}, and any prior P_prior in [0.40, 0.60], the computed P_up SHALL be in (1e-10, 1 - 1e-10).
Validates: Requirements 5.3, 21.5
Property 8: Contradiction is zero when evidence is unidirectional
For any set of cluster LLRs where all clusters have the same sign (all positive or all negative), the computed contradiction score SHALL be 0.0.
Validates: Requirements 7.7, 21.7
Property 9: Contradiction score is bounded in [0, 1]
For any set of cluster LLRs (including mixed positive and negative), the computed contradiction score SHALL be in the range [0.0, 1.0].
Validates: Requirements 7.6
Property 10: Multiplicative confidence is bounded in [0, 1] and suppressed by weak dimensions
For any valid inputs (n_eff_total >= 0, q_values in [0,1], strength in [0,1], regime_confidence_mult in (0,1], contradiction in [0,1], data_quality in [0,1]), the computed confidence SHALL be in [0, 1]. Furthermore, if any single dimension (data_quality, 1-contradiction, or C_quality) is below 0.01, the resulting confidence SHALL be below 0.10.
Validates: Requirements 8.3, 8.5, 21.6
Property 11: Fractional Kelly sizing is bounded and respects negative edge
For any valid inputs (P_up in (0,1), b in [1.2, 3.0], confidence in [0,1], data_quality in [0,1], contradiction in [0,1], max_position_pct > 0), the computed portfolio_pct SHALL be in [0, max_position_pct]. When f_kelly = (P_up * b - (1 - P_up)) / b <= 0, portfolio_pct SHALL be exactly 0.
Validates: Requirements 14.4, 14.7, 21.8, 21.9
Property 12: Posterior state JSON round-trip
For any valid V3Posterior state (p_up, log_odds, strength, confidence, contradiction, n_eff, data_quality, regime), serializing to JSON and deserializing SHALL produce an equivalent state within floating-point tolerance (1e-10).
Validates: Requirements 20.1, 21.10
Property 13: Noisy-OR normalized exposure is bounded in [0, 1]
For any overlap values O_k in [0, 1] for each dimension (geo, supply, commodity, sector) with fixed positive weights, the normalized macro exposure E_macro SHALL be in [0.0, 1.0], reaching exactly 1.0 when all O_k = 1.0.
Validates: Requirements 9.1, 9.2
Property 14: Competitive LLR is clamped to [-1.25, 1.25]
For any source LLR, shrunk correlation (non-negative), graph distance (1-3), and pattern confidence in [0,1], the computed competitive LLR SHALL be in [-1.25, 1.25].
Validates: Requirements 10.4, 10.5
Property 15: Graph attenuation is zero beyond max distance
For any inputs where graph distance > 3, the computed attenuation SHALL be 0.0, producing zero competitive LLR regardless of other parameters.
Validates: Requirements 10.3
Property 16: Projection evidence state decays toward zero
For any initial evidence state A_t and regime decay phi in (0, 1), the projected state A_projected_h = phi^h * A_t SHALL have |A_projected_h| < |A_t| for all h >= 1, converging toward 0 as h increases.
Validates: Requirements 11.1, 11.3
Property 17: Data quality score is bounded in [0, 1]
For any valid inputs (extraction_failure_rate in [0,1], extraction_conf_i in [0,1], impact_i in [0,1], age_newest_hours >= 0, N_valid >= 0, N_source_types >= 0), the computed data_quality_score SHALL be in [0, 1].
Validates: Requirements 17.6
Property 18: Stop loss is below entry price and take profit is above
For any entry_price > 0, stop_distance_pct in [0.005, 1.0), and reward ratio b >= 1.2, the computed stop_loss SHALL be less than entry_price and take_profit SHALL be greater than entry_price.
Validates: Requirements 16.2, 16.3
Property 19: Trailing stop never decreases
For any sequence of current prices and trailing stop computations, each new trailing_stop value SHALL be >= the previous trailing_stop value (monotonically non-decreasing).
Validates: Requirements 16.5
Property 20: Regime classification is exhaustive and deterministic
For any valid market data inputs (closing_prices of sufficient length, daily_returns, ATR_20 > 0), the regime classification SHALL produce exactly one of {panic, trend_following, mean_reversion, uncertainty} and the same inputs SHALL always produce the same classification.
Validates: Requirements 6.2, 6.3, 6.4, 6.5
Property 21: EvidenceUnit normalization preserves field ranges
For any valid company, macro, or competitive signal input, the normalized EvidenceUnit SHALL have: direction in {-1, 0, +1}, sentiment_strength in [0, 1], impact in [0, 1], extraction_conf in [0, 1], source_cred in [0, 1], novelty in [0, 1], and event_base_rate in (0, 1].
Validates: Requirements 1.1, 1.2, 1.3, 1.4, 1.7
Property 22: Portfolio heat rejection is correct
For any set of open positions with stop distances, if the sum of (position_value × stop_distance_pct) exceeds max_portfolio_heat × portfolio_value, then new position entry SHALL be rejected.
Validates: Requirements 15.3, 15.5
Property 23: Tier auto-adjustment obeys downgrade-any, upgrade-all logic
For any TierMetrics, if ANY single downgrade condition is met (profit_factor < 1.0 OR drawdown > 0.12 OR calibration_error > 0.20 OR sharpe < 0), the result SHALL be "downgrade". An "upgrade" SHALL only occur when ALL upgrade conditions are simultaneously met.
Validates: Requirements 18.3, 18.4
Error Handling
Fail-Closed Philosophy
The v3 engine follows a fail-closed design: when in doubt, suppress the trade rather than emit a false signal.
| Error Scenario | Response | Fallback |
|---|---|---|
v3_engine_enabled flag unreadable |
Default to heuristic mode | Log warning |
| Unhandled exception in v3 pipeline | Fall back to heuristic for that cycle | Log ERROR with traceback, record in metadata |
| Missing market data for regime | Default to "uncertainty" regime | Most conservative multipliers |
| Missing source statistics | q_source = 0.0 (neutral prior) | Source treated as untrusted |
| Missing realized_vol_20d | Use default 0.25 annualized | Conservative volatility estimate |
| Division by zero in n_eff | Return n_eff = 1.0 (single signal) | Denominator guard |
| NaN/Inf in any computation | Clamp to boundary, log warning | Never propagate NaN to output |
| data_quality < 0.50 | Force informational mode | Suppress trade recommendation |
| No company evidence (only macro/competitive) | Force informational | Unless macro_only_enabled |
Numerical Guards
All mathematical functions include:
- Sigmoid overflow: Guard
exp(-x)for x > 500 or x < -500 - Log domain: Guard
log(x)with x > 0 check;log2(0)treated as 0 in entropy - Division by zero: All denominators checked > 0 before division
- NaN propagation: All outputs validated with
math.isnan()check before storage - Clamp boundaries: Final values clamped to documented ranges
Graceful Degradation Chain
v3 pipeline error → heuristic fallback → informational mode → no recommendation
Each level preserves audit trail via output metadata.
Testing Strategy
Dual Testing Approach
Property-Based Tests (Hypothesis):
- Library:
hypothesis(already in use in this project) - Configuration:
@settings(max_examples=100)minimum per property - File naming:
tests/test_pbt_v3_*.py - Each property test tagged with:
# Feature: math-core-v3-engine, Property N: <title> - One property-based test per correctness property (23 properties → 23 PBT tests)
Unit Tests (pytest):
- Specific examples with known inputs/outputs for each formula
- Edge cases: zero inputs, boundary values, NaN handling
- Integration between components (e.g., full pipeline from EvidenceUnit to recommendation)
- Error handling paths (DB errors, missing data, feature flag states)
Property Test Organization
| Test File | Properties Covered | Module Under Test |
|---|---|---|
tests/test_pbt_v3_reliability.py |
1, 2, 3, 4, 21 | scoring.py (q_i, p_correct, LLR) |
tests/test_pbt_v3_clustering.py |
5, 6 | worker.py (n_eff, cluster LLR) |
tests/test_pbt_v3_posterior.py |
7, 8, 9, 10, 12, 20 | bayesian.py, contradiction.py, worker.py |
tests/test_pbt_v3_layers.py |
13, 14, 15 | interpolation.py, signal_propagation.py |
tests/test_pbt_v3_projection.py |
16 | projection.py |
tests/test_pbt_v3_decision.py |
11, 17, 18, 19, 22 | position_sizer.py, stop_loss_manager.py, eligibility.py |
tests/test_pbt_v3_tier.py |
23 | risk/engine.py |
Hypothesis Strategy Design
Key custom strategies for generating valid inputs:
from hypothesis import strategies as st
# EvidenceUnit generator
evidence_units = st.builds(
EvidenceUnit,
symbol=st.text(min_size=1, max_size=5),
layer=st.sampled_from(["company", "macro", "competitive"]),
direction=st.sampled_from([-1, 0, 1]),
sentiment_strength=st.floats(min_value=0.0, max_value=1.0),
impact=st.floats(min_value=0.0, max_value=1.0),
extraction_conf=st.floats(min_value=0.0, max_value=1.0),
source_cred=st.floats(min_value=0.0, max_value=1.0),
novelty=st.floats(min_value=0.0, max_value=1.0),
event_base_rate=st.floats(min_value=0.01, max_value=1.0),
...
)
# Cluster LLR list generator
cluster_llrs = st.lists(
st.floats(min_value=-2.5, max_value=2.5),
min_size=1, max_size=20,
)
# Regime generator
regimes = st.sampled_from(["panic", "trend_following", "mean_reversion", "uncertainty"])
Unit Test Coverage
| Area | Key Example Tests |
|---|---|
| EvidenceUnit normalization | Company signal → correct fields; macro → correct horizon mapping |
| q_i pipeline | Known inputs → known outputs for each sub-formula |
| LLR conversion | p_correct=0.60 → LLR≈0.405; direction=-1 → negative LLR |
| Clustering | 3 identical articles → n_eff < 3; independent → n_eff = N |
| Posterior | Empty evidence → P_up=0.50; strong bullish → P_up > 0.60 |
| Contradiction | All bullish → 0; equal split → high score |
| Confidence | Zero data quality → near-zero confidence |
| Macro LLR | Full exposure → max LLR ≈ 1.10; zero overlap → LLR ≈ 0 |
| Competitive | Distance 4 → zero propagation; direct rival → attenuated signal |
| Kelly sizing | p_win=0.3, b=2 → f_kelly < 0 → size = 0 |
| Stops | Entry=100, stop_dist=0.02 → stop=98, TP > 100 |
| Feature flag | Flag false → heuristic path; flag true → v3 path |
| Error fallback | v3 raises → heuristic runs, error logged |
Integration Tests
- Full pipeline: raw signals → EvidenceUnit → q_i → LLR → cluster → posterior → recommendation
- Feature flag toggle: verify clean switch between pipelines mid-run
- JSONB round-trip: store v3 output in PostgreSQL JSONB, retrieve and verify
- Regime transitions: price series that crosses regime boundaries