Defence Q&A — Hard Bits Only
This covers questions the slides don't answer. Everything visible on slides (descriptive stats, charts, procedure steps, KB framework definition, RQs, conclusion, recommendations) is intentionally omitted — those you can read off the slides.
Hard likely from sharp examiners
Trick traps — know the nuance
Note: every answer assumes the examiner has already seen the slide. Don't re-explain what's visible — jump straight to the reasoning behind it.
Why hasn't KB been applied to Japanese before? Hard
Two reasons. First, most KB research originates from Japanese and Korean labs working on English as Foreign Language — English was the natural test domain. Second, Japanese adds complexity English doesn't have: three interwriting scripts, kanji recognition as prerequisite for comprehension, fundamentally different syntax (SOV, particles). You can't simply translate an English KB study design; the reading task is qualitatively different.
Why allow furigana + dictionary? Doesn't that compromise the measure? Hard
This is assisted reading — the goal was not to test kanji recognition in isolation but to see whether KB helps learners connect concepts in a passage they can decode. Beginners relying solely on unassisted kanji would fail at the decoding stage, making the structural comprehension question unanswerable. The assistance narrows the construct to conceptual understanding. However, the MCQ may partially measure decoding ability (scaffolded by answer options) while KB demands independent proposition extraction — explaining the weak KB–gain correlation.
Why use the same passage and same items for pre and post? Trick
Deliberate pilot design choice: same passage + same items allows a clean within-group comparison to detect score direction. It does introduce a practice threat. However, the high-pre group actually declined (mean Δ = −0.33), which argues against a pure practice effect — if everyone just memorised answers from pre-test, high-pre wouldn't go down. For a future controlled study, parallel test forms are recommended.
Why use both paired t-test and Wilcoxon? Hard
Shapiro–Wilk showed all score distributions were non-normal (pre p = 0.019, post p = <0.001, delta p = 0.004) — expected for bounded scores with a ceiling effect. The paired t-test was still used as primary because CLT robustness holds at n ≥ 30 (here n = 31). Wilcoxon serves as the non-parametric supplement. Both agree (t p = 0.017, Wilcoxon p = 0.014) — the convergence strengthens the result. The positive direction is not an artifact of distributional assumptions.
Why report both Cohen's d and Hedges' g? Hard
They measure different things. Cohen's d (paired, within-group) = 0.399 for the pre–post comparison — small-to-medium effect. Hedges' g = 0.988 (large) for the independent median-split comparison. Hedges' g includes a small-sample correction that Cohen's d lacks, making it more appropriate for the split groups (n = 19 vs n = 12). d measures change within individuals; g measures difference between the two groups.
The mean gain is only 0.81 out of 10. Is that meaningful? Hard
The absolute magnitude is small, but context matters. (1) This is a pilot with no control group — the question is direction, not magnitude. (2) The ceiling effect compressed the range: the pre-test mean was already 7.13/10, leaving little room to improve. (3) The effect size (d = 0.399) is small-to-medium — non-trivial for a single-session intervention. Without a control group, we cannot say whether the 0.81 is due to KB, practice, or other factors. That's the whole point of a pilot.
Why is there no correlation between KB score and reading gain? Hard
Spearman ρ = 0.198, p = 0.284 — not significant. The key mechanism: in Japanese, the MCQ and KB tasks tap different constructs. The MCQ provides answer options that scaffold kanji recognition (recognition-based), while KB demands independent proposition extraction from the passage (production-based). KB score thus captures a broader construct — structure + decoding ability — that the MCQ doesn't measure. This is consistent with Alkhateeb (2015) and Funaoi (2011): KB's benefit may lie in long-term retention, which was not measured here. Also, the study is underpowered for a correlation analysis (n = 31).
Why did the high-pre group decline? Trick
Not because KB harmed them. The decline is test-retest noise at the ceiling — when scores are near maximum, random variation is more likely to push scores down than up. The Pearson r = −0.549 between pre-score and gain is a well-known statistical artefact of ceiling effects, not evidence of a real performance drop. This is a measurement issue, not an intervention issue.
How does schema theory explain the ceiling effect specifically? Hard
The passage (わたしのうち, Minna no Nihongo Ch. 10) was confirmed by the lecturer as existing course material. Many students already had well-developed schemas for this content. Schema theory predicts that when learners already possess relevant background knowledge, comprehension is high regardless of the intervention — they're activating existing schemas rather than building new ones. 51.6% scored ≥ 8 at pre-test; the test was likely too easy for the top half of the class. This is one of the clearest lessons from the pilot: future studies need passages that aren't already course material, or need harder items that target inference rather than factual recall.
Participants reported kanji difficulty despite furigana — how? Hard
Furigana provides phonetic reading but doesn't convey meaning on its own — learners still need to connect the phonetic form to the kanji meaning. For true beginners, even with furigana + dictionary, the cognitive load of decoding multiple kanji while also performing the novel KB reconstruction task is substantial. This aligns with Tabata-Sandom (2016): beginners over-rely on pop-up lookups, trapped in bottom-up processing. The bottleneck isn't recognition — it's converting decoded tokens into relational understanding.
Why did engagement drop from pre to post? Is it disinterest? Hard
Engaged participants: 77.1% → 54.3%. The most parsimonious explanation is cognitive fatigue, not disinterest. Participants completed 6 phases in sequence (~30–40 minutes) on top of processing Japanese kanji and performing the novel KB task. This is extraneous cognitive load from task stacking. The TAM results (PU = 3.67, PEoU = 3.55) and feedback ("helps learning" as the #1 theme) argue against disinterest. The fix: split sessions, add breaks, reduce phase count.
Why is the TAM threshold 3.5? Hard
On a 5-point Likert scale, 3.5 is the midpoint between "neutral" (3) and "agree" (4). It's a common threshold in TAM literature — scores above it indicate participants lean toward agreement that the system is useful and easy to use, rather than remaining ambivalent. Both PU (3.67) and PEoU (3.55) clear this threshold, but the margin is modest — suggesting there's room to improve both the perceived usefulness (maybe through better feedback during reconstruction) and ease of use (the initial confusion theme).
What is the "task format gap" you describe? Hard
In English KB studies, the MCQ and KB task measure the same construct: understanding of conceptual connections. In this Japanese study, both had furigana + dictionary, but the MCQ provides answer options that scaffold extraction (recognition-based), while KB requires independent proposition extraction (production-based). KB score captures a broader ability — structure + independent decoding — that the MCQ does not. This construct mismatch explains the weak KB–gain correlation (ρ = 0.198) and is a Japanese-specific artifact that doesn't appear in English KB studies.
Why do you think KB's effect might be in long-term retention rather than immediate gain? Hard
Prior KB research (Alkhateeb 2015, Funaoi 2011) consistently shows KB's benefit emerges in delayed tests, not immediate post-tests. The mechanism: KB forces learners to actively reconstruct the passage's conceptual structure, which builds deeper encoding. This encoding advantage manifests as better retention over time, not necessarily better immediate recall. Since we only measured immediate post-test, we may have missed the window where KB's effect appears. A delayed post-test (2–4 weeks) is needed to confirm — and is the #1 recommendation for the next study.
What was the value of the calibration session? Hard
The Day 2 calibration session (data excluded from analysis) identified 3 ambiguous distractors and 2 items with uncovered vocabulary. These were revised before Day 3, resulting in 10 course-appropriate, unambiguous items. Without this step, those problems would have introduced construct-irrelevant variance into the pre/post-test scores — we'd be measuring item quality, not comprehension. This is one of the procedural contributions from the pilot: calibrate instruments before collecting data.
How does the sample size affect what you can conclude? Hard
n = 31 is adequate for the paired t-test (CLT robustness at n ≥ 30) but underpowered for the correlation analysis and the median split subgroups (n = 19 vs n = 12). Small samples inflate effect size estimates and reduce power to detect real effects. The wide confidence interval on the KB–gain correlation (ρ = 0.198, p = 0.284) likely reflects low power rather than a true null — a moderate correlation might exist but couldn't be detected at this n. Future studies need at least n ≥ 60 for reliable subgroup analyses.
If an examiner asks "So does KB actually work for Japanese?" — how do you answer? Trick
"This pilot shows that KB is feasible and the score direction is positive, but I cannot claim it 'works' in the causal sense from this study alone. The absence of a control group means the improvement could partly reflect practice effects, familiarity, or the passage being course material. What I can say is that the infrastructure works, participants accept the tool, and the direction is promising enough to warrant a controlled experiment — which is exactly what this pilot was designed to justify."
What would you say to an examiner who challenges the novelty? Trick
"The KB framework itself is not novel — Hirashima introduced it in 2015. My contribution is threefold: (1) applying it to a language it has never been applied to — Japanese, which has unique structural challenges; (2) building a modern open-source platform that integrates KB with a complete research workflow; and (3) producing the first documented pilot procedure with calibration data, ready for a controlled experiment. The novelty is in the application, the engineering, and the procedural groundwork — not in re-inventing the framework."
What would you do differently? Hard
(1) Add a control group from the start — even a simple active control (read and summarise). (2) Use parallel test forms for pre and post. (3) Add a delayed post-test at 2–4 weeks. (4) Split the 6 phases across multiple sessions to reduce fatigue. (5) Include a wider proficiency range or harder passages to avoid the ceiling. (6) Track individual kanji lookup frequency to understand the decoding bottleneck precisely. Note: many of these are "do in the next study" items, not things the pilot should have done — the pilot's job was to surface these issues.