Lesson 9 Defense Prep

Numbers & Formulas Reference

Every formula, every number, every calculation chain in your paper — one page

1. Kit-Build Score Formula

Kit-Build Score
S = (M / TGoal) × 100%

Example from your data: a learner with 5 matching propositions out of 7 gets 5/7 × 100 = 71.4%.

A proposition is a meaningful connection between two concepts through a linking word. The goal map for "わたしのうち" encodes 7 propositions covering spatial relationships (house near park, library/coffee shop in front of park) and functional activities (borrowing books, reading, drinking coffee).

Kit-Build Score StatValue
Mean~73%
Median~71%
Min0%
Max100%
SD~25%
Avg reconstructed propositions~5 of 7

2. Reading Comprehension Scores

Both pre-test and post-test used the same 10 multiple-choice questions. Passage: "わたしのうち" (Minna no Nihongo Ch. 10) with furigana. Dictionaries allowed throughout.

MeasureNMeanMedianMinMaxSD
Pre-test317.0683102.41
Post-test317.7782102.34
Delta (Δ)31+0.711-641.99

Mean score increased from 7.06 → 7.77 out of 10. The mean gain was +0.71 points.

Score Trend Breakdown

TrendCount%
Improved (Δ > 0)2064.5%
Same (Δ = 0)619.4%
Declined (Δ < 0)516.1%

Average gain among those who improved: ~1.53 points. Most gains were small (0–2 points), consistent with the ceiling effect.

3. Score Change (Delta)

Score Change
Δ = Post-test Score − Pre-test Score

4. Inferential Statistics Chain

1
Shapiro-Wilk normality test → all distributions non-normal (expected with ceiling effect)
VariableWpVerdict
Pre-test0.9160.019Non-normal
Post-test0.847< 0.001Non-normal
Delta0.8890.004Non-normal
2
Paired t-test (primary) → significant positive score direction
Paired t-test
t = (Mean Difference) / (SDdiff / √n)
3
Wilcoxon signed-rank (nonparametric confirmatory) → confirms
TestStatisticpInterpretation
Paired t-testt = 2.2190.017Significant, d = 0.399
Wilcoxon signed-rankW = 52.0000.014Nonparametric confirmatory
4
Outlier detection (IQR) → 1 outlier (Δ = −6), below lower fence of −3.00
5
Sensitivity analysis → after excluding outlier (n=30): t = 3.520, p < 0.001, d = 0.643 — result robust
Reading the chain: Shapiro-Wilk says non-normal → but CLT + paired design justify t-test → t-test says p = 0.017 (significant) → Wilcoxon confirms p = 0.014 → outlier removal makes it even stronger (p < 0.001). Both parametric and nonparametric converge on the same conclusion.

5. Effect Size Formulas

Cohen's d (paired)
d = Meandiff / SDdiff
Hedges' g (independent groups)
g ≈ (M1 − M2) / SDpooled × correction

6. Correlations

Spearman ρ (Kit-Build Score ↔ Score Change)
ρ = 0.198, p = 0.284
Pearson r (Pre-test Score ↔ Gain)
r = −0.549, p = 0.001
Key distinction: Spearman ρ (Kit-Build ↔ test) is NOT significant. Pearson r (pre-test ↔ gain) IS significant. These answer different questions. The first asks "does reconstruction predict improvement?" The second asks "does baseline level predict improvement?"

7. Ceiling Effect & Median Split

Median Split
Median pre-test = 8.0
GroupnMean PreMean PostMean ΔSD Δ
Low (pre ≤ 8)195.687.21+1.531.78
High (pre > 8)129.429.08−0.331.92

Group difference tests:

TestStatisticpEffect Size
Welch's t-testt = 2.7010.013Hedges' g = 0.988 (large)
Mann-Whitney UU = 168.00.027Nonparametric confirmatory
The ceiling effect story in numbers: Over half the sample scored ≥ 8/10 on the pre-test. The low group gained +1.53 points. The high group lost −0.33 points. The difference is large (g = 0.988) and significant (p = 0.013). This means the test was too easy for advanced beginners, compressing the measurable range of improvement.

8. Technology Acceptance Model

5-point Likert scale (1 = Strongly Disagree, 5 = Strongly Agree). Acceptance threshold = 3.5.

ConstructMeanSDVerdict
Perceived Usefulness (PU)~4.0~0.7Above threshold ✓
Perceived Ease of Use (PEoU)~4.0~0.7Above threshold ✓

All 10 individual items (PU1–PU5, PEoU1–PEoU5) scored above the 3.5 threshold. No construct exhibited uniformly low ratings.

9. Engagement & Data Quality

Completion duration thresholds: < 60s total = rushed (6s per item). Classification:

LevelPre-testPost-test
Engaged27 (87%)19 (61%)
Moderate69
Marginal03
Rushed24

Engagement dropped from 87% → 61% engaged. Interpreted as fatigue from multi-phase online session (cognitive load theory).

Validity Exclusions

StatusCount
Valid (both tests)31
Post-test invalid2
Both invalid2
Total35

10. Qualitative Feedback Summary

ThemeCategoryFreq
Helps learningPositive24
More vocabulary/practiceSuggestion14
Initial confusionDifficulty12
No difficulty reportedDifficulty9
Better UI/tutorialSuggestion6
Simple/easy to usePositive5
Needs clearer instructionsDifficulty7
Kanji difficultyDifficulty4
Audio/pronunciationSuggestion4
More content varietySuggestion4
Interactive/enjoyablePositive3
New/effective methodPositive2

11. The Full Conclusion Chain

How the paper gets from raw data to "positive score direction":

1
35 participants took the study. 4 excluded for rushed responses → n = 31 valid.
2
Pre-test mean = 7.06, post-test mean = 7.77 → Δ = +0.71 (positive direction).
3
Paired t-test: t = 2.219, p = 0.017, d = 0.399 → statistically significant, small-to-medium effect.
4
Wilcoxon confirms: W = 52.000, p = 0.014 → robust to non-normality.
5
Outlier removal (n=30): t = 3.520, p < 0.001, d = 0.643 → not driven by extremes.
6
But: Pearson r = −0.549 (p = 0.001) between pre-test and gain → ceiling effect present.
7
Median split: low group gained +1.53, high group lost −0.33 (g = 0.988, p = 0.013) → ceiling confirmed quantitatively.
8
Spearman ρ = 0.198 (p = 0.284) between Kit-Build score and Δ → no significant link (construct contamination from Kanji).
9
TAM: PU > 3.5, PEoU > 3.5 → positive initial acceptance.
10
Therefore: Yomilink is feasible, shows positive score direction, has initial acceptance — but needs a controlled study to claim effectiveness.
The one number that matters most for defense: The paired t-test gives p = 0.017. This is statistically significant. But your design is a single-group pilot. So you say: "We found a statistically significant positive score direction (p = 0.017, Cohen's d = 0.399), but this is preliminary evidence from a feasibility pilot, not causal evidence of effectiveness."

⚡ Quick Check

1. Your Kit-Build score formula is S = M / T_Goal × 100%. If a learner matches 5 out of 7 propositions, what is their score?

2. Your paired t-test gives p = 0.017. What does this mean in context?

3. The low pre-test group gained +1.53 points and the high group lost −0.33 points. Hedges' g = 0.988. What does g = 0.988 mean?

4. Spearman ρ between Kit-Build score and score change is 0.198 (p = 0.284). What is the best interpretation?

5. What is the single most important number to remember for defense?

Primary Source

Field, A. (2018). Discovering Statistics Using IBM SPSS Statistics (5th ed.). SAGE. Chapters on paired t-tests, effect sizes, and the Central Limit Theorem.

Reminder

Ask me anything that's unclear. If you need me to walk through any formula step-by-step, or explain why a particular number is what it is, just say so.