Comprehensive step-by-step documentation for performing Chi-Square Goodness-of-Fit and Test of Independence on categorical datasets in DATES.
The Chi-Square Test module in DATES evaluates categorical distributions and cross-tabulated contingency tables to assess whether observed frequency counts conform to expected theoretical proportions or whether two categorical variables exhibit statistically significant association.
Supported Chi-Square Modes:
The top toolbar header and sidebar panel allow you to configure test mode, categorical variable selection, significance level ($\alpha$), effect size threshold, and numerical rounding:
| Control / Parameter | Description | Why it is used | When to select / set |
|---|---|---|---|
| Upload Data | Uploads your .csv, .xlsx, or .xls dataset into memory. |
Loads raw observational data and populates categorical column selectors. | At the start of every analysis session. |
| Sheet Selector | Selects the active worksheet from multi-sheet Excel workbooks. | Ensures calculations run on the correct data sheet. | When uploading multi-sheet workbooks. |
| Analysis Mode | Toggles between Test of Independence (two-way cross-tabulation) and Goodness of Fit (single variable). |
Determines whether frequency expectations are computed across two categorical variables or against a single target vector. | Select Independence for two-variable contingency tables; select Goodness of Fit for single categorical distribution checks. |
| Row Variable (Factor A) | Selects the categorical variable representing rows in the contingency table (Independence mode). | Defines row categories for frequency grouping. | Select categorical column for Row Factor A. |
| Column Variable (Factor B) | Selects the categorical variable representing columns in the contingency table (Independence mode). | Defines column categories for frequency grouping. | Select categorical column for Column Factor B. |
| Single Category Variable | Selects the categorical variable for Goodness-of-Fit testing. | Provides observed category frequency counts for single-variable testing. | Select categorical column in Goodness-of-Fit mode. |
| Significance Level (Alpha) | Significance threshold (e.g., 0.05 for 5%, 0.01 for 1%). |
Establishes the critical rejection region boundary for p-values. | Set to 0.05 for standard research or 0.01 for strict confidence requirements. |
| Effect Size Threshold | Sets Cramér's V effect size threshold (Small: w = 0.10, Medium: w = 0.30, Large: w = 0.50). |
Categorizes the magnitude of association between categorical factors. | Select threshold based on scientific domain reporting standards. |
| Decimals | Controls rounding precision (1 to 6 decimal places) in output tables. | Formats frequency tables and test statistics for journal reporting. | Adjust based on required numerical precision. |
DATES accepts dataset spreadsheets in standard .xlsx, .xls, or .csv formats. Data should contain categorical text labels or discrete coded factor columns:
| Subject_ID | Treatment_Factor | Response_Outcome | Severity_Grade |
|---|---|---|---|
| Obs-001 | Method_A | Positive | Moderate |
| Obs-002 | Method_A | Positive | Mild |
| Obs-003 | Method_B | Negative | Severe |
| Obs-004 | Method_B | Positive | Moderate |
| Obs-005 | Method_A | Negative | Mild |
| Obs-006 | Method_B | Positive | Moderate |
The Chi-Square test compares observed frequency counts (O) in each cell against expected frequency counts (E) under the null hypothesis of independence or equal distribution. Below are the plain text definitions:
Formula Description:
Chi-Square = Sum of ((Observed Frequency - Expected Frequency)^2 / Expected Frequency) across all table cells.
Formula Description:
Expected Cell Frequency = (Row Total * Column Total) / Grand Total
Degrees of Freedom: df = (Number of Rows - 1) * (Number of Columns - 1)
Formula Description:
Expected Category Frequency = Grand Total / Number of Categories (under equal distribution null)
Degrees of Freedom: df = Number of Categories - 1
Formula Description:
Cramér's V = Square Root of (Chi-Square / (Grand Total * Minimum of (Rows - 1, Columns - 1)))
Measures the strength of association on a scale from 0.0 (no association) to 1.0 (perfect association).
.csv or .xlsx file.Test of Independence for two-variable contingency tables or Goodness of Fit for single-variable distributions..xlsx), Word documents (.docx), PowerPoint presentations (.pptx), or publication-grade PNG images.Below is an example of an output summary table generated for a Chi-Square Test of Independence between two categorical factors:
| Comparison / Variables | Chi-Square Statistic | df | p-Value | Cramér's V | Effect Magnitude | Association Status |
|---|---|---|---|---|---|---|
| Treatment_Factor vs Response_Outcome | 12.450 | 2 | 0.0020 | 0.342 | Medium Association | Reject H0 (Significant Association) |
| Treatment_Factor vs Severity_Grade | 1.180 | 2 | 0.5543 | 0.082 | Negligible | Fail to Reject H0 (Independent) |
v < 0.10 (Negligible), 0.10 to 0.30 (Small), 0.30 to 0.50 (Medium), and v > 0.50 (Large).Chi-Square tests require that no more than 20% of expected cell frequencies are less than 5, and all expected cell frequencies exceed 1. If small sample sizes lead to low expected frequencies, consider collapsing adjacent categories or running Fisher's Exact Test.
Ensure that all observational units in your dataset are independent (each subject contributes to exactly one cell in the contingency table). For repeated categorical measurements, use McNemar's test instead of standard Chi-Square.
If you use the DATES Chi-Square Test module in published scientific research, please cite it as follows: