Module III · Method Validation Case 01Módulo III · Validación de Métodos Caso 01

Road Accident System — Method Evidence Report

Sistema de Accidentes de Tránsito — Method Evidence Report

Application of the generic Method Validation Handbook to the official Module II ESS and the 4,999-record Road Accident dataset.

Aplicación del Method Validation Handbook genérico al ESS oficial del Módulo II y al dataset real de 4,999 accidentes.

Execution verdictVeredicto de ejecución
EXECUTED WITH FLAGS
Indices and diagnostics were calculated. Scientific acceptance remains unevaluated until Module IV. Se calcularon índices y diagnósticos. La aceptación científica permanece sin evaluar hasta el Módulo IV.

1. Black-box execution rule

1. Regla de ejecución de caja negra

The ESS is frozen. This execution preserves the ordinal target, selected variables, coordinate pair, separate casualty objective, and deferred causality.El ESS está fijado. Esta ejecución preserva el objetivo ordinal, las variables seleccionadas, el par de coordenadas, el objetivo separado de víctimas y la causalidad diferida.
Decision boundary: no metric below is labeled scientifically sufficient or insufficient. Module III reports the measurements and flags. Module IV decides.Frontera decisional: ningún índice se declara científicamente suficiente o insuficiente. El Módulo III reporta mediciones y banderas. El Módulo IV decide.

2. Data snapshot and target integrity

2. Snapshot de datos e integridad del objetivo

4,999recorded accidentsaccidentes registrados
4,103Slight · 82.1%
563Serious · 11.3%
333Fatal · 6.7%
Source correction: the label “Fetal” was recoded to “Fatal” for execution. The source value remains a formal data-quality defect.Corrección de fuente: “Fetal” fue recodificado a “Fatal” para ejecutar. El valor original sigue siendo un defecto formal.

3. Conditional-variable quality gates

3. Puertas de calidad de variables condicionales

VariableObserved structureEstructura observadaExecution resultResultado de ejecución
Junction_Control970 / 4,999 = 19.4% “Data missing or out of range”RESTRICTED — not advanced into core modeling
Carriageway_Hazards4,961 / 4,999 = 99.2% “None”SPARSE — conditional evidence highly unstable
Speed_limit4,942 / 4,999 = 98.9% at 30DESCRIPTIVE ONLY, as required by ESS
Urban_or_Rural_Area1 unique state: UrbanNOT APPLICABLE for comparison
Police_Force2 levels; 94.1% Metropolitan PoliceADMINISTRATIVE CONTROL ONLY

4. Association and dependence execution

4. Ejecución de asociación y dependencia

Categorical predictors were evaluated against the ordinal severity classes using contingency structure, permutation-calibrated chi-square, bias-corrected Cramér’s V, and normalized mutual information.

Los predictores categóricos se evaluaron contra las clases ordinales mediante contingencia, chi-cuadrado calibrado por permutación, V de Cramér corregida e información mutua normalizada.

VariableLevelsχ²Permutation pCramér’s VNMIExpected cells <5
Local_Authority_(District)15353.010.0030.1800.040421/45
Weather_Conditions834.250.0100.0450.005212/24
Junction_Detail932.070.0270.0400.00376/27
Road_Type523.970.0070.0400.00393/15
Road_Surface_Conditions411.400.0900.0230.00184/12
Light_Conditions55.320.6780.0000.00117/15
Vehicle_Type1321.590.6150.0000.00299/39
Day_of_Week711.280.4880.0000.00100/21
Measured pattern: Local_Authority_(District) produced the largest global categorical association (V = 0.180333861472649237267518174121505580842494964599609375000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000). Weather, Road_Type, and Junction_Detail produced smaller signals. Several tables contain sparse expected cells, so permutation evidence is shown instead of relying only on asymptotic p-values.Patrón medido: Local_Authority_(District) produjo la mayor asociación global (V = 0.180333861472649237267518174121505580842494964599609375000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000000). Weather, Road_Type y Junction_Detail produjeron señales menores. Varias tablas poseen celdas esperadas escasas; por eso se muestra permutación.

5. Information execution

5. Ejecución informacional

Highest normalized MILocal_Authority_(District): 0.0404
Environmental NMIWeather: 0.0052 · Surface: 0.0018 · Light: 0.0011
Infrastructure NMIRoad Type: 0.0039 · Junction: 0.0037
Vehicle NMIVehicle Type: 0.0029
All normalized MI values are small on this coding and sample. Module III reports that fact without deciding whether “small” is acceptable for the case.Todos los valores de NMI son pequeños bajo esta codificación y muestra. El Módulo III reporta el hecho sin decidir si “pequeño” es aceptable.

6. Time execution

6. Ejecución temporal

6.1 Hour-of-day segmentation

6.1 Segmentación por hora del día

Hour binSlightSeriousFatalSerious/Fatal rate
00–05266582824.4%
06–097661276620.1%
10–15132415111016.5%
16–1912401439316.0%
20–23507843619.1%

χ² = 27.66 · p = 0.0005 · Cramér’s V = 0.044 · NMI = 0.0029

6.2 Daily aggregated severity series

6.2 Serie diaria agregada de severidad

364observed calendar daysdías observados
0.508ACF lag 1
0.412ACF lag 7
500.2Ljung–Box Q(7), p<0.001
Temporal boundary: these are aggregated event-record patterns. They do not measure accident risk over time because traffic exposure and non-accident denominators are absent.Frontera temporal: son patrones agregados de eventos registrados. No miden riesgo de accidente porque faltan exposición y denominadores sin accidente.

7. Spatial execution

7. Ejecución espacial

0.075Moran’s I · 8-nearest neighbors
0.0033permutation p-valuep por permutación
4,409unique coordinate pairspares únicos
459duplicated coordinate pairspares duplicados
A positive global spatial autocorrelation signal was measured for ordinal severity scores. The analysis describes clustering among recorded accident events only.Se midió una señal positiva de autocorrelación espacial global para la severidad ordinal. El análisis describe agrupamiento entre accidentes registrados solamente.

8. Cross-objective execution: severity and casualties

8. Ejecución entre objetivos: severidad y víctimas

Spearman ρ0.0474p = 0.0008
Kruskal–Wallis H17.133p = 0.0002
Cramér’s V (casualties 1 / 2 / 3+)0.0375NMI = 0.0035
The relationship is measurable but weak in magnitude. Number_of_Casualties remains a separate consequence objective and was not merged into the severity model.La relación es medible pero débil en magnitud. Number_of_Casualties permanece como objetivo separado y no se mezcló con el modelo de severidad.

9. Prediction execution — chronological holdout

9. Ejecución predictiva — holdout cronológico

The last 20% of accidents by Accident Date was reserved as an out-of-time holdout. The predictor set follows the ESS core variables. Junction_Control, Speed_limit, Urban_or_Rural_Area, Accident_Index, casualties, and post-outcome fields were excluded.

El último 20% por Accident Date se reservó como holdout fuera del tiempo. El conjunto predictor sigue las variables centrales del ESS. Se excluyeron Junction_Control, Speed_limit, Urban_or_Rural_Area, Accident_Index, víctimas y campos posteriores.

MethodAccuracyBalanced Acc.Macro F1Macro ROC-AUCLog LossRecall Slight / Serious / Fatal
Majority baseline0.8760.3330.3110.5000.4651.000 / 0.000 / 0.000
Multinomial Logistic0.3930.2670.2380.4481.1280.425 / 0.158 / 0.217
Random Forest0.5550.2820.2720.4700.8820.623 / 0.050 / 0.174
Critical generalization flag: both fitted models performed worse than the majority baseline on several out-of-time metrics, especially log loss and balanced accuracy. Fatal and Serious recall remained unstable. This is a measured execution result, not yet a Module IV rejection decision.Bandera crítica de generalización: ambos modelos rindieron peor que la línea base mayoritaria en varias métricas fuera del tiempo, especialmente log loss y balanced accuracy. El recall de Fatal y Serious fue inestable. Es un resultado medido, no todavía una decisión de rechazo.

10. Prediction stability — stratified cross-validation

10. Estabilidad predictiva — validación estratificada

MethodBalanced AccuracyMacro F1Macro ROC-AUC
Multinomial Logistic0.478 ± 0.0180.340 ± 0.0130.633 ± 0.013
Random Forest0.457 ± 0.0190.385 ± 0.0170.654 ± 0.010
Random stratified folds produce materially better scores than the chronological holdout. This disagreement is itself an important stability and temporal-shift signal for Module IV.Los folds estratificados aleatorios producen mejores métricas que el holdout cronológico. Esa discrepancia es una señal importante de estabilidad y cambio temporal para el Módulo IV.

11. Explanation gate

11. Puerta de explicación

STATUS: NOT ACTIVATED AS FINAL EVIDENCE. The ESS requires explanation only after predictive approval. Because prediction produced critical generalization flags, SHAP/LIME/ALE explanations would currently explain unstable model behavior rather than an approved model.ESTADO: NO ACTIVADA COMO EVIDENCIA FINAL. El ESS exige explicación después de aprobación predictiva. Como la predicción produjo banderas críticas, SHAP/LIME/ALE explicarían comportamiento inestable y no un modelo aprobado.

12. Causality gate

12. Puerta causal

DEFERRED — NOT EXECUTED.
No treatment, intervention assignment, counterfactual comparison, or identification design exists in the ESS. No causal algorithm was selected or run.No existe tratamiento, asignación, comparación contrafactual ni diseño de identificación. No se seleccionó ni ejecutó algoritmo causal.

13. Integrated evidence matrix

13. Matriz integrada de evidencia

Evidence familyExecution statusMeasured technical signalFlag for Module IV
AssociationEXECUTEDDistrict strongest; several smaller categorical signalsSparse cells and small magnitudes
DependenceEXECUTEDPermutation-supported dependence for District, Weather, Road Type, JunctionObservational only
InformationEXECUTEDAll normalized MI values smallEstimator and sparse-category sensitivity
TimeEXECUTED WITH FLAGSHour segmentation and daily autocorrelation detectedNo exposure denominator; aggregation effects
SpaceEXECUTED WITH FLAGSMoran’s I = 0.075, permutation p = 0.0033Recorded-event geography only
PredictionEXECUTED WITH CRITICAL FLAGSChronological generalization weak; random CV materially higherTemporal shift and rare-class failure
ExplanationGATE CLOSEDAwaiting predictive decisionDo not explain unapproved model
CausalityDEFERREDNot executedNo causal architecture
IntegrationEXECUTEDNative results preserved without collapsing into one scoreContradictions retained

14. Failure and robustness registry

14. Registro de fallos y robustez

FlagStatusEvidence
Source-label defectOPEN“Fetal” required recoding to “Fatal”
Sparse-category riskOPENWeather, Light, Vehicle, Surface and Junction tables contain expected cells below 5
Conditional-variable failureOPENJunction_Control 19.4% invalid; Carriageway_Hazards 99.2% None
Temporal stabilityOPENRandom CV and chronological holdout disagree materially
Rare-class predictionCRITICALFatal recall on chronological holdout: Logistic 0.217; RF 0.174
Spatial supportCONDITIONAL459 coordinate pairs duplicated; no traffic exposure denominator
ReproducibilityDOCUMENTEDRandom seed 42; explicit methods and split architecture recorded

15. Method Evidence Report handoff

15. Entrega del Method Evidence Report

METHOD EVIDENCE REPORT — CASE 01

ESS: KESM-ESS-M2-C01-v2
TARGET: Accident_Severity (ordinal)
RECORDS: 4,999
EXECUTION STATUS: EXECUTED WITH FLAGS

ASSOCIATION / DEPENDENCE: EXECUTED
INFORMATION: EXECUTED
TIME: EXECUTED WITH LIMITS
SPACE: EXECUTED WITH LIMITS
PREDICTION: EXECUTED WITH CRITICAL GENERALIZATION FLAGS
EXPLANATION: NOT ACTIVATED
CAUSALITY: DEFERRED
INTEGRATION: EXECUTED

MODULE IV DECISION:
NOT EVALUATED IN MODULE III

16. What Module IV must decide

16. Lo que debe decidir el Módulo IV

  • Whether the categorical association magnitudes are scientifically relevant despite statistical evidence.
  • Si las magnitudes categóricas son científicamente relevantes pese a la evidencia estadística.
  • Whether the temporal and spatial signals are usable under the no-exposure boundary.
  • Si las señales temporales y espaciales son utilizables bajo la frontera sin exposición.
  • Whether the prediction architecture is rejected, revised, or conditionally continued.
  • Si la arquitectura predictiva se rechaza, revisa o continúa condicionalmente.
  • Whether explanation may be activated after a new predictive iteration.
  • Si la explicación puede activarse después de una nueva iteración predictiva.

17. Case closing

17. Cierre del caso

The handbook worked exactly as designed: it did not force success. It exposed measurable association, information, temporal and spatial structure, while also exposing sparse support, temporal instability, weak out-of-time prediction, and rare-class failure.El manual funcionó exactamente como fue diseñado: no forzó éxito. Expuso asociación, información y estructura temporal/espacial medibles, y también soporte escaso, inestabilidad temporal, predicción fuera del tiempo débil y fallo de clase rara.