BI4Humans Business Investigation Framework v1.0

Case #2 — Police Incident Operational IntelligenceCaso #2 — Inteligencia Operacional de Incidentes Policiales

A complete evidence-driven investigation using Databricks, Apache Spark and PySpark, with every script located exactly where it answers the business question.

Una investigación completa guiada por evidencia utilizando Databricks, Apache Spark y PySpark, con cada script ubicado exactamente donde responde la pregunta de negocio.

Default ENES / ENDatabricksApache SparkGoverned PipelineDecision Intelligence

Executive SnapshotResumen Ejecutivo

786,074Incident recordsRegistros de incidentes
29.04%Larceny Theft shareParticipación de Larceny Theft
+53.71%Burglary change, 2019→2020Cambio de Burglary, 2019→2020

Official Data SourceFuente Oficial de Datos

The case uses the public dataset Police Department Incident Reports: 2018 to Present, published through DataSF by the San Francisco Police Department, and loaded into the governed Databricks table workspace.default.police_incidents.

El caso utiliza el dataset público Police Department Incident Reports: 2018 to Present, publicado mediante DataSF por el San Francisco Police Department, y cargado en la tabla gobernada de Databricks workspace.default.police_incidents.

The findings represent the specific public extract available in the workspace, not the complete live operational police database.

Los hallazgos representan el extracto público específico disponible en el workspace, no la base operacional policial completa y en vivo.

Databricks + Apache Spark + Governed Pipeline

Official Public Source ↓ Secure Ingestion ↓ Bronze — Raw and Traceable ↓ Silver — Validated and Standardized ↓ Gold — Evidence, KPIs and Decision Models ↓ Databricks Governance + Apache Spark Parallel Processing ↓ Executive Reports, Dashboards and Controlled Decisions

Databricks provides the governed collaborative environment; Apache Spark provides distributed parallel processing; PySpark expresses the investigation as reusable code. In production, the pipeline would include lineage, role-based access, least privilege, encryption, monitoring, privacy controls and separate development, test and production environments.

Databricks proporciona el ambiente colaborativo gobernado; Apache Spark proporciona procesamiento paralelo distribuido; PySpark expresa la investigación como código reutilizable. En producción, el pipeline incluiría linaje, acceso basado en roles, mínimo privilegio, cifrado, monitoreo, controles de privacidad y ambientes separados de desarrollo, prueba y producción.

Investigation TreeÁrbol de Investigación

Governed Table ✓ ├── Scope & Completeness ✓ ├── Annual Baseline ✓ ├── Category Distribution ✓ │ └── Larceny Trend ✓ ├── 2019 vs 2020 Comparison ✓ │ ├── Burglary Branch ✓ │ │ ├── District Concentration ✓ │ │ ├── Neighborhood Concentration ✓ │ │ ├── Day-of-Week Hypothesis ✗ │ │ ├── Hour-of-Day Pattern ✓ │ │ └── Neighborhood × Hour Branch ◌ │ └── Motor Vehicle Theft — Future Branch └── Operational Priority Model — Future Branch

NavigationNavegación

Investigation #1

Load the Governed Analytical TableCargar la Tabla Analítica Gobernada

CompletedConfidence: HighFoundation
Question → Hypothesis → Spark Investigation → Evidence → Pattern → Decision → Next Question

1. Business QuestionPregunta de Negocio

Can the investigation begin from an identified, governed and reproducible table?¿Puede comenzar la investigación desde una tabla identificada, gobernada y reproducible?

2. Why This MattersPor Qué Importa

A professional analysis must establish exactly which object is being analyzed before any metric is calculated.Un análisis profesional debe establecer exactamente qué objeto se analiza antes de calcular cualquier métrica.

3. Working HypothesisHipótesis de Trabajo

The Databricks workspace contains a prepared incident table suitable for analysis.El workspace de Databricks contiene una tabla preparada de incidentes adecuada para el análisis.

4. Spark Investigation

from pyspark.sql.functions import *

df = spark.table("workspace.default.police_incidents")

record_count = df.count()
column_count = len(df.columns)

print(f"Records: {record_count:,}")
print(f"Columns: {column_count}")
df.printSchema()
display(df.limit(10))

5. Actual EvidenceEvidencia Real

786,074 records and 29 analytical fields were available in the governed table.Había 786,074 registros y 29 campos analíticos disponibles en la tabla gobernada.

6. Pattern ObservedPatrón Observado

The table is large enough to demonstrate distributed processing while remaining practical for an instructional investigation.La tabla es suficientemente grande para demostrar procesamiento distribuido y a la vez práctica para una investigación educativa.

7. Executive InterpretationInterpretación Ejecutiva

The analytical object is identifiable, repeatable and ready for controlled investigation.El objeto analítico es identificable, repetible y está listo para una investigación controlada.

8. Decision ImpactImpacto en la Decisión

Authorize the first analytical cycle using this governed table as the single source for the case.Autorizar el primer ciclo analítico utilizando esta tabla gobernada como fuente única del caso.

9. Next InvestigationPróxima Investigación

Validate temporal scope and completeness.Validar el alcance temporal y la integridad.

Investigation #2

Validate Time Coverage and Partial-Year RiskValidar Cobertura Temporal y Riesgo de Año Parcial

CompletedConfidence: HighFoundation
Question → Hypothesis → Spark Investigation → Evidence → Pattern → Decision → Next Question

1. Business QuestionPregunta de Negocio

What time period does the published extract actually cover?¿Qué período cubre realmente el extracto publicado?

2. Why This MattersPor Qué Importa

Comparing incomplete years with complete years can create false operational conclusions.Comparar años incompletos con años completos puede producir conclusiones operacionales falsas.

3. Working HypothesisHipótesis de Trabajo

The table covers 2018–2024, but 2024 is only partially represented.La tabla cubre 2018–2024, pero 2024 está representado solo parcialmente.

4. Spark Investigation

df_dates = (
    df.withColumn("Incident_Date_Parsed", to_date("Incident_Date", "yyyy/MM/dd"))
)

date_scope = df_dates.agg(
    min("Incident_Date_Parsed").alias("Minimum_Date"),
    max("Incident_Date_Parsed").alias("Maximum_Date")
)

yearly_counts = (
    df.groupBy("Incident_Year")
      .agg(count("Incident_ID").alias("Total_Incidents"))
      .orderBy("Incident_Year")
)

display(date_scope)
display(yearly_counts)

5. Actual EvidenceEvidencia Real

The observed period runs from January 1, 2018 through March 12, 2024. The 2024 count is 19,273 and represents a partial year.El período observado va del 1 de enero de 2018 al 12 de marzo de 2024. El conteo de 2024 es 19,273 y representa un año parcial.

6. Pattern ObservedPatrón Observado

2018–2023 are broadly comparable full-year periods; 2024 requires explicit exclusion or normalization.2018–2023 son períodos anuales completos ampliamente comparables; 2024 requiere exclusión explícita o normalización.

7. Executive InterpretationInterpretación Ejecutiva

The investigation must separate complete and incomplete periods before any trend ranking.La investigación debe separar períodos completos e incompletos antes de cualquier ranking de tendencias.

8. Decision ImpactImpacto en la Decisión

Use 2018–2023 for full-year comparisons and label 2024 as partial.Usar 2018–2023 para comparaciones anuales completas y etiquetar 2024 como parcial.

9. Next InvestigationPróxima Investigación

Establish the annual incident baseline.Establecer la línea base anual de incidentes.

Investigation #3

Establish the Annual Incident BaselineEstablecer la Línea Base Anual de Incidentes

CompletedConfidence: HighFoundation
Question → Hypothesis → Spark Investigation → Evidence → Pattern → Decision → Next Question

1. Business QuestionPregunta de Negocio

How did total incident volume evolve by year?¿Cómo evolucionó el volumen total de incidentes por año?

2. Why This MattersPor Qué Importa

The annual baseline reveals structural shifts that require category-level decomposition.La línea base anual revela cambios estructurales que requieren descomposición por categoría.

3. Working HypothesisHipótesis de Trabajo

Total incidents fell sharply in 2020 and then partially recovered.Los incidentes totales cayeron fuertemente en 2020 y luego se recuperaron parcialmente.

4. Spark Investigation

annual_incidents = (
    df.groupBy("Incident_Year")
      .agg(count("Incident_ID").alias("Total_Incidents"))
      .orderBy("Incident_Year")
)

display(annual_incidents)

5. Actual EvidenceEvidencia Real

2018: 143,630; 2019: 138,330; 2020: 112,069; 2021: 121,583; 2022: 127,205; 2023: 123,984; 2024 partial: 19,273.2018: 143,630; 2019: 138,330; 2020: 112,069; 2021: 121,583; 2022: 127,205; 2023: 123,984; 2024 parcial: 19,273.

6. Pattern ObservedPatrón Observado

The largest complete-year disruption occurred in 2020, followed by recovery without a full return to the 2018 level.La mayor ruptura entre años completos ocurrió en 2020, seguida de recuperación sin retorno completo al nivel de 2018.

7. Executive InterpretationInterpretación Ejecutiva

A global decline does not reveal which operational categories changed or in what direction.Una disminución global no revela qué categorías operacionales cambiaron ni en qué dirección.

8. Decision ImpactImpacto en la Decisión

Open a category-distribution branch before interpreting the 2020 decline.Abrir una rama de distribución por categorías antes de interpretar la caída de 2020.

9. Next InvestigationPróxima Investigación

Measure category volume and percentage share.Medir volumen y participación porcentual por categoría.

Investigation #4

Measure Incident Category DistributionMedir la Distribución de Categorías de Incidentes

CompletedConfidence: HighIntermediate
Question → Hypothesis → Spark Investigation → Evidence → Pattern → Decision → Next Question

1. Business QuestionPregunta de Negocio

Which incident categories dominate the published operational workload?¿Qué categorías dominan la carga operacional publicada?

2. Why This MattersPor Qué Importa

Counts without percentage share do not show relative operational importance.Los conteos sin participación porcentual no muestran la importancia operacional relativa.

3. Working HypothesisHipótesis de Trabajo

Larceny Theft represents the largest single category.Larceny Theft representa la categoría individual más grande.

4. Spark Investigation

total_incidents = df.count()

category_share = (
    df.groupBy("Incident_Category")
      .agg(count("Incident_ID").alias("Incidents"))
      .withColumn(
          "Share_Percentage",
          round((col("Incidents") / lit(total_incidents)) * 100, 2)
      )
      .orderBy(col("Incidents").desc())
)

display(category_share)

5. Actual EvidenceEvidencia Real

Larceny Theft: 228,239 incidents, 29.04%. Burglary: 45,784 incidents, 5.82%. Motor Vehicle Theft: 44,421 incidents, 5.65%.Larceny Theft: 228,239 incidentes, 29.04%. Burglary: 45,784 incidentes, 5.82%. Motor Vehicle Theft: 44,421 incidentes, 5.65%.

6. Pattern ObservedPatrón Observado

One category alone accounts for almost three of every ten published incidents.Una sola categoría representa casi tres de cada diez incidentes publicados.

7. Executive InterpretationInterpretación Ejecutiva

Larceny Theft is the natural first branch, but category concentration alone does not explain the 2020 change.Larceny Theft es la primera rama natural, pero la concentración por categoría no explica por sí sola el cambio de 2020.

8. Decision ImpactImpacto en la Decisión

Analyze Larceny Theft by year and normalize it against annual totals.Analizar Larceny Theft por año y normalizarlo contra los totales anuales.

9. Next InvestigationPróxima Investigación

Build the Larceny Theft annual trend.Construir la tendencia anual de Larceny Theft.

Investigation #5

Normalize the Larceny Theft TrendNormalizar la Tendencia de Larceny Theft

CompletedConfidence: HighIntermediate
Question → Hypothesis → Spark Investigation → Evidence → Pattern → Decision → Next Question

1. Business QuestionPregunta de Negocio

Did Larceny Theft fall only in count, or also as a share of all incidents?¿Larceny Theft cayó solo en conteo o también como participación de todos los incidentes?

2. Why This MattersPor Qué Importa

A category can decline in count simply because the entire dataset declined. Normalization tests whether its operational weight also changed.Una categoría puede disminuir en conteo simplemente porque disminuyó todo el dataset. La normalización prueba si también cambió su peso operacional.

3. Working HypothesisHipótesis de Trabajo

Larceny Theft declined in both volume and share during 2020.Larceny Theft disminuyó tanto en volumen como en participación durante 2020.

4. Spark Investigation

annual_total = (
    df.groupBy("Incident_Year")
      .agg(count("Incident_ID").alias("Total_Incidents"))
)

larceny_year = (
    df.filter(col("Incident_Category") == "Larceny Theft")
      .groupBy("Incident_Year")
      .agg(count("Incident_ID").alias("Larceny_Incidents"))
)

larceny_trend = (
    larceny_year.join(annual_total, "Incident_Year")
      .withColumn(
          "Larceny_Share",
          round((col("Larceny_Incidents") / col("Total_Incidents")) * 100, 2)
      )
      .orderBy("Incident_Year")
)

display(larceny_trend)

5. Actual EvidenceEvidencia Real

2019 share: 31.82%. 2020 share: 25.16%. The count fell from 44,020 to 28,196.Participación 2019: 31.82%. Participación 2020: 25.16%. El conteo cayó de 44,020 a 28,196.

6. Pattern ObservedPatrón Observado

The decline was real in both absolute and relative terms.La disminución fue real tanto en términos absolutos como relativos.

7. Executive InterpretationInterpretación Ejecutiva

Larceny Theft explains much of the global decline, but it cannot support the claim that every category declined.Larceny Theft explica gran parte de la caída global, pero no respalda la afirmación de que todas las categorías disminuyeron.

8. Decision ImpactImpacto en la Decisión

Compare every category between 2019 and 2020.Comparar todas las categorías entre 2019 y 2020.

9. Next InvestigationPróxima Investigación

Create a category pivot and rank absolute and percentage change.Crear un pivot por categoría y clasificar cambio absoluto y porcentual.

Investigation #6

Compare Every Category: 2019 vs 2020Comparar Todas las Categorías: 2019 vs 2020

CompletedConfidence: HighAdvanced
Question → Hypothesis → Spark Investigation → Evidence → Pattern → Decision → Next Question

1. Business QuestionPregunta de Negocio

Did all incident categories decline during 2020?¿Disminuyeron todas las categorías de incidentes durante 2020?

2. Why This MattersPor Qué Importa

A global trend can hide opposing movements among its components.Una tendencia global puede ocultar movimientos opuestos entre sus componentes.

3. Working HypothesisHipótesis de Trabajo

Most categories declined, but some operational categories increased.La mayoría de las categorías disminuyeron, pero algunas categorías operacionales aumentaron.

4. Spark Investigation

category_comparison = (
    df.filter(col("Incident_Year").isin(2019, 2020))
      .groupBy("Incident_Category")
      .pivot("Incident_Year", [2019, 2020])
      .agg(count("Incident_ID"))
      .fillna(0)
      .withColumn("Difference", col("2020") - col("2019"))
      .withColumn("Absolute_Change", abs(col("Difference")))
      .withColumn(
          "Percent_Change",
          when(
              col("2019") > 0,
              round((col("Difference") / col("2019")) * 100, 2)
          )
      )
      .orderBy(col("Absolute_Change").desc())
)

display(category_comparison)

5. Actual EvidenceEvidencia Real

Larceny Theft −35.95%; Burglary +53.71%; Motor Vehicle Theft +40.94%.Larceny Theft −35.95%; Burglary +53.71%; Motor Vehicle Theft +40.94%.

6. Pattern ObservedPatrón Observado

The 2020 disruption changed the incident mix rather than simply reducing all categories uniformly.La ruptura de 2020 cambió la composición de incidentes en lugar de reducir uniformemente todas las categorías.

7. Executive InterpretationInterpretación Ejecutiva

The hypothesis that every category declined is rejected. Burglary becomes the strongest explanatory branch because it combines high percentage growth with substantial absolute impact.Se rechaza la hipótesis de que todas las categorías disminuyeron. Burglary se convierte en la rama explicativa más fuerte porque combina alto crecimiento porcentual con impacto absoluto sustancial.

8. Decision ImpactImpacto en la Decisión

Prioritize Burglary for district-level decomposition.Priorizar Burglary para una descomposición a nivel de distrito.

9. Next InvestigationPróxima Investigación

Measure which districts contributed most to the increase.Medir qué distritos contribuyeron más al aumento.

Investigation #7

Decompose Burglary Growth by Police DistrictDescomponer el Crecimiento de Burglary por Distrito Policial

CompletedConfidence: HighAdvanced
Question → Hypothesis → Spark Investigation → Evidence → Pattern → Decision → Next Question

1. Business QuestionPregunta de Negocio

Which police districts contributed most to the burglary increase?¿Qué distritos policiales contribuyeron más al aumento de burglary?

2. Why This MattersPor Qué Importa

Operational resources are assigned geographically; citywide averages are not directly actionable.Los recursos operacionales se asignan geográficamente; los promedios de toda la ciudad no son directamente accionables.

3. Working HypothesisHipótesis de Trabajo

Burglary growth was concentrated in a limited number of districts.El crecimiento de burglary estuvo concentrado en un número limitado de distritos.

4. Spark Investigation

burglary_district = (
    df.filter(
        (col("Incident_Category") == "Burglary") &
        (col("Incident_Year").isin(2019, 2020))
    )
    .groupBy("Police_District")
    .pivot("Incident_Year", [2019, 2020])
    .agg(count("Incident_ID"))
    .fillna(0)
    .withColumn("Difference", col("2020") - col("2019"))
    .withColumn(
        "Percent_Change",
        when(col("2019") > 0, round((col("2020")-col("2019"))/col("2019")*100,2))
    )
    .orderBy(col("Difference").desc())
)

display(burglary_district)

5. Actual EvidenceEvidencia Real

Northern +832 (+82.54%); Mission +463 (+80.38%); Park +415 (+125.76%); Richmond +350 (+92.11%).Northern +832 (+82.54%); Mission +463 (+80.38%); Park +415 (+125.76%); Richmond +350 (+92.11%).

6. Pattern ObservedPatrón Observado

Northern had the greatest absolute impact, while Park had the greatest relative increase.Northern tuvo el mayor impacto absoluto, mientras Park tuvo el mayor aumento relativo.

7. Executive InterpretationInterpretación Ejecutiva

Absolute and relative rankings answer different management questions. Both must be preserved.Los rankings absoluto y relativo responden preguntas gerenciales diferentes. Ambos deben conservarse.

8. Decision ImpactImpacto en la Decisión

Open a Northern District neighborhood branch because it contributed the largest number of additional incidents.Abrir una rama por barrios de Northern District porque contribuyó el mayor número de incidentes adicionales.

9. Next InvestigationPróxima Investigación

Measure neighborhood concentration within Northern District.Medir la concentración por barrios dentro de Northern District.

Investigation #8

Measure Neighborhood Concentration in Northern DistrictMedir la Concentración por Barrios en Northern District

CompletedConfidence: HighAdvanced
Question → Hypothesis → Spark Investigation → Evidence → Pattern → Decision → Next Question

1. Business QuestionPregunta de Negocio

Was Northern District burglary growth broadly distributed or concentrated in a few neighborhoods?¿El crecimiento de burglary en Northern District estuvo ampliamente distribuido o concentrado en pocos barrios?

2. Why This MattersPor Qué Importa

Concentration reveals whether broad deployment or targeted intervention is more appropriate.La concentración revela si es más apropiado un despliegue amplio o una intervención focalizada.

3. Working HypothesisHipótesis de Trabajo

A small number of neighborhoods explain most Northern District burglary incidents in 2020.Un pequeño número de barrios explica la mayoría de los incidentes de burglary de Northern District en 2020.

4. Spark Investigation

northern_2020 = (
    df.filter(
        (col("Incident_Category") == "Burglary") &
        (col("Incident_Year") == 2020) &
        (col("Police_District") == "Northern")
    )
)

total_northern = northern_2020.count()

neighborhood_share = (
    northern_2020.groupBy("Analysis_Neighborhood")
      .agg(count("Incident_ID").alias("Burglary_Incidents"))
      .withColumn(
          "Share_Percentage",
          round((col("Burglary_Incidents") / lit(total_northern)) * 100, 2)
      )
      .orderBy(col("Burglary_Incidents").desc())
)

display(neighborhood_share)

5. Actual EvidenceEvidencia Real

Marina 25.65%; Hayes Valley 19.46%; Pacific Heights 18.86%. Combined: 63.97%.Marina 25.65%; Hayes Valley 19.46%; Pacific Heights 18.86%. Combinados: 63.97%.

6. Pattern ObservedPatrón Observado

Nearly two thirds of Northern District burglary incidents were concentrated in three neighborhoods.Casi dos tercios de los incidentes de burglary de Northern District se concentraron en tres barrios.

7. Executive InterpretationInterpretación Ejecutiva

The evidence favors targeted geographic analysis over district-wide generalization.La evidencia favorece un análisis geográfico focalizado sobre una generalización de todo el distrito.

8. Decision ImpactImpacto en la Decisión

Prioritize temporal analysis within the concentrated neighborhoods.Priorizar el análisis temporal dentro de los barrios concentrados.

9. Next InvestigationPróxima Investigación

Test whether day of week provides a strong operational signal.Probar si el día de la semana proporciona una señal operacional fuerte.

Investigation #9

Test the Day-of-Week HypothesisProbar la Hipótesis por Día de Semana

RejectedConfidence: HighIntermediate
Question → Hypothesis → Spark Investigation → Evidence → Pattern → Decision → Next Question

1. Business QuestionPregunta de Negocio

Is burglary strongly concentrated on weekends?¿Burglary está fuertemente concentrado en fines de semana?

2. Why This MattersPor Qué Importa

A strong day pattern could influence shift planning and preventive coverage.Un patrón fuerte por día podría influir en la planificación de turnos y cobertura preventiva.

3. Working HypothesisHipótesis de Trabajo

Weekend days have clearly higher burglary shares.Los días de fin de semana tienen participaciones claramente mayores de burglary.

4. Spark Investigation

burglary_day = (
    df.filter(col("Incident_Category") == "Burglary")
      .groupBy("Incident_Day_of_Week")
      .agg(count("Incident_ID").alias("Burglary_Incidents"))
)

total_burglary = burglary_day.agg(
    sum("Burglary_Incidents").alias("Total")
).first()["Total"]

burglary_day_share = (
    burglary_day.withColumn(
        "Share_Percentage",
        round((col("Burglary_Incidents") / lit(total_burglary)) * 100, 2)
    )
    .orderBy(col("Burglary_Incidents").desc())
)

display(burglary_day_share)

5. Actual EvidenceEvidencia Real

Friday was highest at 15.73%, while Saturday and Sunday were the two lowest shares.Friday fue el más alto con 15.73%, mientras Saturday y Sunday tuvieron las dos participaciones más bajas.

6. Pattern ObservedPatrón Observado

The distribution is relatively even and does not show meaningful weekend concentration.La distribución es relativamente uniforme y no muestra concentración significativa en fines de semana.

7. Executive InterpretationInterpretación Ejecutiva

Day of week is a weak discriminator for resource allocation by itself.El día de la semana es un discriminador débil para asignación de recursos por sí solo.

8. Decision ImpactImpacto en la Decisión

Reject the weekend hypothesis and investigate hour of day.Rechazar la hipótesis de fin de semana e investigar la hora del día.

9. Next InvestigationPróxima Investigación

Extract and analyze incident hour.Extraer y analizar la hora del incidente.

Investigation #10

Analyze Hour-of-Day BehaviorAnalizar el Comportamiento por Hora del Día

CompletedConfidence: Medium-HighAdvanced
Question → Hypothesis → Spark Investigation → Evidence → Pattern → Decision → Next Question

1. Business QuestionPregunta de Negocio

Does hour of day reveal stronger burglary behavior than day of week?¿La hora del día revela un comportamiento de burglary más fuerte que el día de la semana?

2. Why This MattersPor Qué Importa

Hourly profiles can support more precise operational windows than daily averages.Los perfiles horarios pueden apoyar ventanas operacionales más precisas que los promedios diarios.

3. Working HypothesisHipótesis de Trabajo

Burglary shows visible overnight and late-day peaks.Burglary muestra picos visibles durante la madrugada y al final del día.

4. Spark Investigation

df_hour = (
    df.withColumn(
        "Incident_Timestamp",
        to_timestamp("Incident_Datetime", "yyyy/MM/dd hh:mm:ss a")
    )
    .withColumn("Incident_Hour", hour("Incident_Timestamp"))
)

burglary_hour = (
    df_hour.filter(col("Incident_Category") == "Burglary")
      .groupBy("Incident_Hour")
      .agg(count("Incident_ID").alias("Burglary_Incidents"))
      .orderBy("Incident_Hour")
)

display(burglary_hour)

5. Actual EvidenceEvidencia Real

The visible profile shows notable activity around 03:00–05:00 and renewed growth during the late afternoon.El perfil visible muestra actividad notable alrededor de 03:00–05:00 y crecimiento renovado al final de la tarde.

6. Pattern ObservedPatrón Observado

There appears to be more than one operational pattern rather than a single universal peak.Parece existir más de un patrón operacional en lugar de un único pico universal.

7. Executive InterpretationInterpretación Ejecutiva

A citywide hourly profile can still hide different neighborhood behaviors.Un perfil horario de toda la ciudad todavía puede ocultar comportamientos diferentes por barrio.

8. Decision ImpactImpacto en la Decisión

Create a neighborhood × hour behavioral matrix.Crear una matriz de comportamiento barrio × hora.

9. Next InvestigationPróxima Investigación

Compare neighborhood-specific hourly profiles.Comparar perfiles horarios específicos por barrio.

Investigation #11

Build the Neighborhood × Hour Behavioral MatrixConstruir la Matriz de Comportamiento Barrio × Hora

Open BranchConfidence: MediumAdvanced
Question → Hypothesis → Spark Investigation → Evidence → Pattern → Decision → Next Question

1. Business QuestionPregunta de Negocio

Do different neighborhoods exhibit distinct burglary hour profiles?¿Diferentes barrios presentan perfiles horarios distintos de burglary?

2. Why This MattersPor Qué Importa

Distinct profiles would support differentiated operating strategies instead of one citywide schedule.Perfiles distintos apoyarían estrategias operacionales diferenciadas en lugar de un solo horario para toda la ciudad.

3. Working HypothesisHipótesis de Trabajo

Commercial, residential and mixed-use neighborhoods may display different hourly behavior.Los barrios comerciales, residenciales y de uso mixto pueden mostrar comportamientos horarios diferentes.

4. Spark Investigation

hour_neighborhood = (
    df_hour.filter(col("Incident_Category") == "Burglary")
      .groupBy("Analysis_Neighborhood", "Incident_Hour")
      .agg(count("Incident_ID").alias("Burglary_Incidents"))
      .orderBy("Analysis_Neighborhood", "Incident_Hour")
)

display(hour_neighborhood)

financial_profile = (
    hour_neighborhood.filter(
        col("Analysis_Neighborhood") == "Financial District/South Beach"
    )
)

display(financial_profile)

5. Actual EvidenceEvidencia Real

Financial District/South Beach showed substantial overnight activity, including elevated counts around 02:00–04:00.Financial District/South Beach mostró actividad sustancial durante la madrugada, incluyendo conteos elevados alrededor de 02:00–04:00.

6. Pattern ObservedPatrón Observado

The visible profile supports the possibility of neighborhood-specific behavior, but external land-use validation is still required.El perfil visible respalda la posibilidad de comportamientos específicos por barrio, pero todavía se requiere validación externa del uso del suelo.

7. Executive InterpretationInterpretación Ejecutiva

This branch is promising but not yet fully closed. It should remain in the future investigation roadmap.Esta rama es prometedora pero aún no está completamente cerrada. Debe permanecer en la hoja de ruta de futuras investigaciones.

8. Decision ImpactImpacto en la Decisión

Preserve the branch for comparative segmentation rather than making causal claims.Conservar la rama para segmentación comparativa en lugar de hacer afirmaciones causales.

9. Next InvestigationPróxima Investigación

Validate missing neighborhood values before deeper segmentation.Validar valores faltantes de barrio antes de una segmentación más profunda.

Investigation #12

Validate Missing Neighborhood ValuesValidar Valores Faltantes de Barrio

RejectedConfidence: HighIntermediate
Question → Hypothesis → Spark Investigation → Evidence → Pattern → Decision → Next Question

1. Business QuestionPregunta de Negocio

Could missing neighborhood values materially distort the burglary analysis?¿Podrían los valores faltantes de barrio distorsionar materialmente el análisis de burglary?

2. Why This MattersPor Qué Importa

Visible NULL rows can appear important without measuring their actual volume.Las filas NULL visibles pueden parecer importantes sin medir su volumen real.

3. Working HypothesisHipótesis de Trabajo

Missing neighborhood values represent a meaningful data-quality limitation.Los valores faltantes de barrio representan una limitación significativa de calidad de datos.

4. Spark Investigation

missing_burglary_neighborhood = (
    df.filter(
        (col("Incident_Category") == "Burglary") &
        (col("Analysis_Neighborhood").isNull())
    )
    .agg(count("Incident_ID").alias("Missing_Neighborhood_Records"))
)

display(missing_burglary_neighborhood)

5. Actual EvidenceEvidencia Real

Only 4 burglary records had a missing Analysis_Neighborhood value.Solo 4 registros de burglary tenían un valor faltante de Analysis_Neighborhood.

6. Pattern ObservedPatrón Observado

The apparent issue was visually noticeable but operationally negligible.El problema aparente era visualmente notable pero operacionalmente insignificante.

7. Executive InterpretationInterpretación Ejecutiva

The data-quality hypothesis is rejected. The investigation can continue without material bias from this field.Se rechaza la hipótesis de calidad de datos. La investigación puede continuar sin sesgo material proveniente de este campo.

8. Decision ImpactImpacto en la Decisión

Close the first information cycle and document future branches.Cerrar el primer ciclo de información y documentar las ramas futuras.

9. Next InvestigationPróxima Investigación

Move toward operational priority modeling, hotspot persistence and governed forecasting.Avanzar hacia modelado de prioridad operacional, persistencia de hotspots y pronóstico gobernado.

Executive Findings and Closed Information CycleHallazgos Ejecutivos y Ciclo de Información Cerrado

Cycle closure: The case defines its scope, gathers sufficient evidence, rejects unsupported hypotheses, closes one coherent information cycle and preserves future branches without pretending that every possible question has been answered.

Cierre del ciclo: El caso define su alcance, reúne evidencia suficiente, rechaza hipótesis no respaldadas, cierra un ciclo coherente de información y conserva ramas futuras sin pretender que todas las preguntas posibles han sido respondidas.

Common MistakesErrores Comunes

Future Investigation RoadmapHoja de Ruta de Investigaciones Futuras

Final Methodological StatementDeclaración Metodológica Final

Databricks is not used merely to run code. It is used as a governed analytical environment where public data moves through a secure, automated and scalable pipeline, while Apache Spark supplies the parallel processing capacity required to transform large operational datasets into reproducible evidence and executive decisions.

Databricks no se utiliza solamente para ejecutar código. Se utiliza como un ambiente analítico gobernado donde los datos públicos se mueven mediante un pipeline seguro, automatizado y escalable, mientras Apache Spark aporta la capacidad de procesamiento paralelo necesaria para transformar grandes datasets operacionales en evidencia reproducible y decisiones ejecutivas.

Cybersecurity and Production NoticeAviso de Ciberseguridad y Producción

This material uses public, non-production data for education and analytical demonstration. Production deployment requires formal legal, privacy, governance, security, performance and operational approval.

Este material utiliza datos públicos y no productivos para educación y demostración analítica. El despliegue en producción requiere aprobación formal legal, de privacidad, gobernanza, seguridad, rendimiento y operaciones.