1. Black-box execution rule
1. Regla de ejecución de caja negra
The ESS is frozen. This execution preserves the ordinal target, selected variables, coordinate pair, separate casualty objective, and deferred causality.El ESS está fijado. Esta ejecución preserva el objetivo ordinal, las variables seleccionadas, el par de coordenadas, el objetivo separado de víctimas y la causalidad diferida.
Decision boundary: no metric below is labeled scientifically sufficient or insufficient. Module III reports the measurements and flags. Module IV decides.Frontera decisional: ningún índice se declara científicamente suficiente o insuficiente. El Módulo III reporta mediciones y banderas. El Módulo IV decide.
2. Data snapshot and target integrity
2. Snapshot de datos e integridad del objetivo
4,999recorded accidentsaccidentes registrados
4,103Slight · 82.1%
563Serious · 11.3%
333Fatal · 6.7%
Source correction: the label “Fetal” was recoded to “Fatal” for execution. The source value remains a formal data-quality defect.Corrección de fuente: “Fetal” fue recodificado a “Fatal” para ejecutar. El valor original sigue siendo un defecto formal.
3. Conditional-variable quality gates
3. Puertas de calidad de variables condicionales
| Variable | Observed structure | Estructura observada | Execution result | Resultado de ejecución |
| Junction_Control | 970 / 4,999 = 19.4% “Data missing or out of range” | RESTRICTED — not advanced into core modeling |
| Carriageway_Hazards | 4,961 / 4,999 = 99.2% “None” | SPARSE — conditional evidence highly unstable |
| Speed_limit | 4,942 / 4,999 = 98.9% at 30 | DESCRIPTIVE ONLY, as required by ESS |
| Urban_or_Rural_Area | 1 unique state: Urban | NOT APPLICABLE for comparison |
| Police_Force | 2 levels; 94.1% Metropolitan Police | ADMINISTRATIVE CONTROL ONLY |
4. Association and dependence execution
4. Ejecución de asociación y dependencia
Categorical predictors were evaluated against the ordinal severity classes using contingency structure, permutation-calibrated chi-square, bias-corrected Cramér’s V, and normalized mutual information.
Los predictores categóricos se evaluaron contra las clases ordinales mediante contingencia, chi-cuadrado calibrado por permutación, V de Cramér corregida e información mutua normalizada.
| Variable | Levels | χ² | Permutation p | Cramér’s V | NMI | Expected cells <5 |
| Local_Authority_(District) | 15 | 353.01 | 0.003 | 0.180 | 0.0404 | 21/45 |
| Weather_Conditions | 8 | 34.25 | 0.010 | 0.045 | 0.0052 | 12/24 |
| Junction_Detail | 9 | 32.07 | 0.027 | 0.040 | 0.0037 | 6/27 |
| Road_Type | 5 | 23.97 | 0.007 | 0.040 | 0.0039 | 3/15 |
| Road_Surface_Conditions | 4 | 11.40 | 0.090 | 0.023 | 0.0018 | 4/12 |
| Light_Conditions | 5 | 5.32 | 0.678 | 0.000 | 0.0011 | 7/15 |
| Vehicle_Type | 13 | 21.59 | 0.615 | 0.000 | 0.0029 | 9/39 |
| Day_of_Week | 7 | 11.28 | 0.488 | 0.000 | 0.0010 | 0/21 |
Measured pattern: Local_Authority_(District) produced the largest global categorical association (V = 0.180333861472649237267518174121505580842494964599609375000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000). Weather, Road_Type, and Junction_Detail produced smaller signals. Several tables contain sparse expected cells, so permutation evidence is shown instead of relying only on asymptotic p-values.Patrón medido: Local_Authority_(District) produjo la mayor asociación global (V = 0.180333861472649237267518174121505580842494964599609375000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000). Weather, Road_Type y Junction_Detail produjeron señales menores. Varias tablas poseen celdas esperadas escasas; por eso se muestra permutación.
5. Information execution
5. Ejecución informacional
Highest normalized MILocal_Authority_(District): 0.0404
Environmental NMIWeather: 0.0052 · Surface: 0.0018 · Light: 0.0011
Infrastructure NMIRoad Type: 0.0039 · Junction: 0.0037
Vehicle NMIVehicle Type: 0.0029
All normalized MI values are small on this coding and sample. Module III reports that fact without deciding whether “small” is acceptable for the case.Todos los valores de NMI son pequeños bajo esta codificación y muestra. El Módulo III reporta el hecho sin decidir si “pequeño” es aceptable.
6. Time execution
6. Ejecución temporal
6.1 Hour-of-day segmentation
6.1 Segmentación por hora del día
| Hour bin | Slight | Serious | Fatal | Serious/Fatal rate |
|---|
| 00–05 | 266 | 58 | 28 | 24.4% |
| 06–09 | 766 | 127 | 66 | 20.1% |
| 10–15 | 1324 | 151 | 110 | 16.5% |
| 16–19 | 1240 | 143 | 93 | 16.0% |
| 20–23 | 507 | 84 | 36 | 19.1% |
χ² = 27.66 · p = 0.0005 · Cramér’s V = 0.044 · NMI = 0.0029
6.2 Daily aggregated severity series
6.2 Serie diaria agregada de severidad
364observed calendar daysdías observados
0.508ACF lag 1
0.412ACF lag 7
500.2Ljung–Box Q(7), p<0.001
Temporal boundary: these are aggregated event-record patterns. They do not measure accident risk over time because traffic exposure and non-accident denominators are absent.Frontera temporal: son patrones agregados de eventos registrados. No miden riesgo de accidente porque faltan exposición y denominadores sin accidente.
7. Spatial execution
7. Ejecución espacial
0.075Moran’s I · 8-nearest neighbors
0.0033permutation p-valuep por permutación
4,409unique coordinate pairspares únicos
459duplicated coordinate pairspares duplicados
A positive global spatial autocorrelation signal was measured for ordinal severity scores. The analysis describes clustering among recorded accident events only.Se midió una señal positiva de autocorrelación espacial global para la severidad ordinal. El análisis describe agrupamiento entre accidentes registrados solamente.
8. Cross-objective execution: severity and casualties
8. Ejecución entre objetivos: severidad y víctimas
| Spearman ρ | 0.0474 | p = 0.0008 |
| Kruskal–Wallis H | 17.133 | p = 0.0002 |
| Cramér’s V (casualties 1 / 2 / 3+) | 0.0375 | NMI = 0.0035 |
The relationship is measurable but weak in magnitude. Number_of_Casualties remains a separate consequence objective and was not merged into the severity model.La relación es medible pero débil en magnitud. Number_of_Casualties permanece como objetivo separado y no se mezcló con el modelo de severidad.
9. Prediction execution — chronological holdout
9. Ejecución predictiva — holdout cronológico
The last 20% of accidents by Accident Date was reserved as an out-of-time holdout. The predictor set follows the ESS core variables. Junction_Control, Speed_limit, Urban_or_Rural_Area, Accident_Index, casualties, and post-outcome fields were excluded.
El último 20% por Accident Date se reservó como holdout fuera del tiempo. El conjunto predictor sigue las variables centrales del ESS. Se excluyeron Junction_Control, Speed_limit, Urban_or_Rural_Area, Accident_Index, víctimas y campos posteriores.
| Method | Accuracy | Balanced Acc. | Macro F1 | Macro ROC-AUC | Log Loss | Recall Slight / Serious / Fatal |
|---|
| Majority baseline | 0.876 | 0.333 | 0.311 | 0.500 | 0.465 | 1.000 / 0.000 / 0.000 |
| Multinomial Logistic | 0.393 | 0.267 | 0.238 | 0.448 | 1.128 | 0.425 / 0.158 / 0.217 |
| Random Forest | 0.555 | 0.282 | 0.272 | 0.470 | 0.882 | 0.623 / 0.050 / 0.174 |
Critical generalization flag: both fitted models performed worse than the majority baseline on several out-of-time metrics, especially log loss and balanced accuracy. Fatal and Serious recall remained unstable. This is a measured execution result, not yet a Module IV rejection decision.Bandera crítica de generalización: ambos modelos rindieron peor que la línea base mayoritaria en varias métricas fuera del tiempo, especialmente log loss y balanced accuracy. El recall de Fatal y Serious fue inestable. Es un resultado medido, no todavía una decisión de rechazo.
10. Prediction stability — stratified cross-validation
10. Estabilidad predictiva — validación estratificada
| Method | Balanced Accuracy | Macro F1 | Macro ROC-AUC |
|---|
| Multinomial Logistic | 0.478 ± 0.018 | 0.340 ± 0.013 | 0.633 ± 0.013 |
| Random Forest | 0.457 ± 0.019 | 0.385 ± 0.017 | 0.654 ± 0.010 |
Random stratified folds produce materially better scores than the chronological holdout. This disagreement is itself an important stability and temporal-shift signal for Module IV.Los folds estratificados aleatorios producen mejores métricas que el holdout cronológico. Esa discrepancia es una señal importante de estabilidad y cambio temporal para el Módulo IV.
11. Explanation gate
11. Puerta de explicación
STATUS: NOT ACTIVATED AS FINAL EVIDENCE. The ESS requires explanation only after predictive approval. Because prediction produced critical generalization flags, SHAP/LIME/ALE explanations would currently explain unstable model behavior rather than an approved model.ESTADO: NO ACTIVADA COMO EVIDENCIA FINAL. El ESS exige explicación después de aprobación predictiva. Como la predicción produjo banderas críticas, SHAP/LIME/ALE explicarían comportamiento inestable y no un modelo aprobado.
12. Causality gate
12. Puerta causal
DEFERRED — NOT EXECUTED.
No treatment, intervention assignment, counterfactual comparison, or identification design exists in the ESS. No causal algorithm was selected or run.No existe tratamiento, asignación, comparación contrafactual ni diseño de identificación. No se seleccionó ni ejecutó algoritmo causal.
13. Integrated evidence matrix
13. Matriz integrada de evidencia
| Evidence family | Execution status | Measured technical signal | Flag for Module IV |
| Association | EXECUTED | District strongest; several smaller categorical signals | Sparse cells and small magnitudes |
| Dependence | EXECUTED | Permutation-supported dependence for District, Weather, Road Type, Junction | Observational only |
| Information | EXECUTED | All normalized MI values small | Estimator and sparse-category sensitivity |
| Time | EXECUTED WITH FLAGS | Hour segmentation and daily autocorrelation detected | No exposure denominator; aggregation effects |
| Space | EXECUTED WITH FLAGS | Moran’s I = 0.075, permutation p = 0.0033 | Recorded-event geography only |
| Prediction | EXECUTED WITH CRITICAL FLAGS | Chronological generalization weak; random CV materially higher | Temporal shift and rare-class failure |
| Explanation | GATE CLOSED | Awaiting predictive decision | Do not explain unapproved model |
| Causality | DEFERRED | Not executed | No causal architecture |
| Integration | EXECUTED | Native results preserved without collapsing into one score | Contradictions retained |
14. Failure and robustness registry
14. Registro de fallos y robustez
| Flag | Status | Evidence |
| Source-label defect | OPEN | “Fetal” required recoding to “Fatal” |
| Sparse-category risk | OPEN | Weather, Light, Vehicle, Surface and Junction tables contain expected cells below 5 |
| Conditional-variable failure | OPEN | Junction_Control 19.4% invalid; Carriageway_Hazards 99.2% None |
| Temporal stability | OPEN | Random CV and chronological holdout disagree materially |
| Rare-class prediction | CRITICAL | Fatal recall on chronological holdout: Logistic 0.217; RF 0.174 |
| Spatial support | CONDITIONAL | 459 coordinate pairs duplicated; no traffic exposure denominator |
| Reproducibility | DOCUMENTED | Random seed 42; explicit methods and split architecture recorded |
15. Method Evidence Report handoff
15. Entrega del Method Evidence Report
METHOD EVIDENCE REPORT — CASE 01
ESS: KESM-ESS-M2-C01-v2
TARGET: Accident_Severity (ordinal)
RECORDS: 4,999
EXECUTION STATUS: EXECUTED WITH FLAGS
ASSOCIATION / DEPENDENCE: EXECUTED
INFORMATION: EXECUTED
TIME: EXECUTED WITH LIMITS
SPACE: EXECUTED WITH LIMITS
PREDICTION: EXECUTED WITH CRITICAL GENERALIZATION FLAGS
EXPLANATION: NOT ACTIVATED
CAUSALITY: DEFERRED
INTEGRATION: EXECUTED
MODULE IV DECISION:
NOT EVALUATED IN MODULE III
16. What Module IV must decide
16. Lo que debe decidir el Módulo IV
- Whether the categorical association magnitudes are scientifically relevant despite statistical evidence.
- Si las magnitudes categóricas son científicamente relevantes pese a la evidencia estadística.
- Whether the temporal and spatial signals are usable under the no-exposure boundary.
- Si las señales temporales y espaciales son utilizables bajo la frontera sin exposición.
- Whether the prediction architecture is rejected, revised, or conditionally continued.
- Si la arquitectura predictiva se rechaza, revisa o continúa condicionalmente.
- Whether explanation may be activated after a new predictive iteration.
- Si la explicación puede activarse después de una nueva iteración predictiva.
17. Case closing
17. Cierre del caso
The handbook worked exactly as designed: it did not force success. It exposed measurable association, information, temporal and spatial structure, while also exposing sparse support, temporal instability, weak out-of-time prediction, and rare-class failure.El manual funcionó exactamente como fue diseñado: no forzó éxito. Expuso asociación, información y estructura temporal/espacial medibles, y también soporte escaso, inestabilidad temporal, predicción fuera del tiempo débil y fallo de clase rara.