Pedagogical Friction Framework

Interactive Quantitative Learning Guide

A tool to understand and explain the quantitative strand of the methodology.

Proposal stage. No participant data have been collected and no findings are reported. Every number here is generated demonstration data. The survey is not designed to produce population estimates or to validate the framework as a psychometric scale.

Toggle "Explain it Simply" Mode

1. Data Cleaning & Missing Data

The survey incorporates role-specific branching. Responses of "Don't know" on infrastructural items are treated as missing during scale construction but are reported as a separate category in descriptive analysis.

The "Don't know" rate is interpreted as substantive evidence rather than only a data quality issue.

The Analogy: Imagine asking students if there's a school policy on using calculators. If 30% say "I don't know", that's not just missing data—it's a massive finding! It means the school has terrible communication.

In our study, if teachers don't know the AI policy, that uncertainty is actual evidence of infrastructural friction.

Insight: 15% missing data. This indicates moderate infrastructural friction.

2. Composite Scoring: Mean vs. Sum

Composite scores utilize mean scoring rather than sum scoring. This ensures scores remain interpretable on the original 5-point Likert metric and transparently handles item-level missingness.

The Analogy: If a student takes 4 out of 5 quizzes, a Sum Score gives them a 0 on the missed quiz, unfairly dropping their total grade. A Mean Score averages the 4 they took.

By taking the average (Mean), the final score stays between 1 and 5, which is much easier to explain than an arbitrary sum like "17 out of 25".

Toggle missing responses to see how the scores react.

Q1:
Q2:
Q3:
Q4:

Sum Score

-
Drops heavily if a question is missed.

Mean Score (Preferred)

-
Stays stable on the 1-5 scale.

3. Item Behaviour in This Sample

Item-total correlations, inter-item correlations, and internal consistency estimates are reported only as preliminary evidence of how items behave within this sample — not as validation of the framework as a scale. The survey is not designed to produce population estimates or to validate the framework psychometrically. Early item statistics inform interpretation and future refinement rather than mid-collection validation.

The Choir Analogy: This asks a narrow question: "Did the people who answered tend to answer these related questions in similar ways?"

If they did, the questions probably belong together for this group of respondents. That is useful to know and worth reporting. What it does not tell us is whether the framework itself is correct, or whether these questions would behave the same way with different educators.

Experiment with the Choir

Adjust how strongly related items move together, and how many items each domain has.

Mean inter-item r

0.00

Mean item-total r

0.00

4. Role Comparisons and Cautions

Significance testing (t-tests, ANOVA) will be conducted to examine group differences. Effect sizes (Cohen’s d, η²) will be reported alongside p-values, as significance alone is insufficient for interpreting findings.

The Analogy: A p-value tells us if a difference is REAL (not a fluke). An Effect Size tells us if the difference MATTERS.

Imagine a weight loss pill. A study might find it causes 0.1 lbs of weight loss, and the p-value says this is statistically significant (real). But an effect size would show that 0.1 lbs is practically meaningless! We need both to tell the full story.

Scenario: Written AI Policy vs. No Policy

We ran a t-test on Existential Friction scores.

p-value: 0.034 *
Cohen's d: 0.62 (Medium)
How to explain this: "Teachers with a written AI policy had lower friction. We know this wasn't a fluke (p < .05), and more importantly, Cohen's d shows that having a policy makes a moderately large, practical difference in their daily experience."