Lesson 11
Defense Prep
Slide-by-Slide Talking Script
Mapped to the actual Typst deck — 28 slides, ~18–20 min. Every slide, what's on it, and exactly what you say.
Your slides are vague on purpose. That's good. Vague slides + confident speaker = examiner focus on YOU, not the screen. The slide is a visual anchor. You are the content. This script fills the gap.
Time Budget
Cover 0:20
Intro 2:30
Framework 1:45
Method 2:45
Results 5:30
Discussion 2:45
Conclusion 1:30
Close 0:15
Total: ~18 min. Target 15–20. Anything over 20 and examiners get restless.
Table of Contents
- Cover — Slide 1
- Intro — Slides 2–4 (problem, Kit-Build, RQs)
- Framework — Slides 5–6 (KB mechanics, related work)
- Methodology — Slides 7–10 (design, Yomilink, procedure, analysis)
- Results — Slides 11–21 (descriptive, RQ1, RQ2, RQ3, inferential, summary)
- Discussion — Slides 22–25 (claims, interpretation, gap, limitations)
- Conclusion — Slides 26–27 (conclusion, recommendations)
- Closing — Slide 28
How to use this script: Don't memorise word-for-word. Read each slide's "what to say" section. Practice out loud 3 times. By the third run, the structure will be in your bones and you'll sound natural.
COVER — 1 slide, ~20 seconds
On the slide: Full thesis title, your name, NIM, "A Preliminary Pilot Study"
What to say
"Good morning. I'm Dicha Zelianivan Arkana, NIM 2241720002. Today I'm presenting my thesis — a preliminary pilot study on applying Kit-Build concept mapping for Japanese reading comprehension among beginner learners."
Don't read the title. They can read. Just greet, introduce yourself, and state what the study is about in your own words. Hit the word "pilot" early — it sets expectations.
INTRO — 3 slides, ~2:30
→ Transition: "Let me start with why this matters."
On the slide: Three scripts, SOV order, ~709k learners in Indonesia, memorisation ≠ connecting ideas. Right column: "The gap — learners decode words but miss how concepts connect."
What to say
"Japanese has three writing systems — Hiragana, Katakana, and Kanji. For Indonesian learners whose native language uses Latin script, this is a substantial barrier. Add to that the SOV word order, verb conjugations, and particles — all structurally different from Indonesian — and you get a real comprehension problem. Indonesia has over 700,000 Japanese learners — second most in the world — so motivation is high. But motivation doesn't remove the structural difficulty. The core gap is this: learners can decode individual words, especially with furigana help, but they struggle to connect those words into meaningful relationships. They see the pieces, not the picture."
Pause after "second most in the world." Let that land. It justifies why this research matters in Indonesia specifically.
→ Transition: "So how do we help learners see the connections? One approach is concept mapping."
On the slide: Concept mapping = nodes + links. Open-ended hard to assess. Kit-Build: teacher → goal map, system → kit, learner → learner map, system diagnoses matching/missing/excessive. Score formula: S = (M / T_Goal) × 100%.
What to say
"Concept mapping is a technique where learners draw nodes for concepts and links for relationships. It externalises understanding. The problem: open-ended maps are impossible to assess automatically — every student draws something different. Kit-Build, developed by Hirashima in 2015, solves this. The teacher creates a goal map — the correct understanding. The system decomposes it into components — nodes and links — forming a 'kit.' Learners reconstruct the map from the kit. Then the system automatically diagnoses matching, missing, and excessive propositions at the link level. The score is the proportion of correctly reconstructed propositions — matching links divided by total goal links, times one hundred percent."
Key move: When you explain the formula — S = M / T × 100% — point to it on the slide. This is one of the few things worth reading aloud because it's the core metric.
→ Transition: "Kit-Build has been studied, but there's a gap — which brings us to my research questions."
On the slide: Three RQs (score direction, KB reconstruction performance, learner perceptions). Right column: scope — beginner learners, 2 Marketing classes, assisted reading, one-group pretest-posttest pilot, no causal claims.
What to say
"I investigated three questions. First: what is the within-group reading comprehension score direction after the Kit-Build activity? Note the word 'direction' — not 'improvement' or 'effect.' Second: how do beginner Japanese learners perform on Kit-Build reconstruction tasks? Third: what are learner perceptions and experiences? The scope is deliberately narrow. This is a pilot study with beginner Japanese learners from two Marketing Management classes at Polinema. Reading was assisted — furigana and dictionary were available. The design is one-group pretest-posttest. I want to be very clear upfront: this study makes no causal claims and no claim of platform superiority."
Critical: Say "no causal claims" and "no claim of platform superiority" verbatim. Front-loading these caveats disarms examiners who would otherwise challenge you on them later. It signals that you know exactly what your study is and isn't.
FRAMEWORK — 2 slides, ~1:45
→ Transition: "Let me walk through the Kit-Build framework in more detail."
On the slide: Four boxes in a vertical flowchart: Goal Map → Kit → Learner Map → Diagnosis (Matching/Missing/Excessive). Caption: "Close-ended → automatic proposition diagnosis → formative feedback."
What to say
"Here is the full workflow. Step one: the teacher prepares a goal map — the target understanding of the reading passage. Step two: the system decomposes the goal map into a kit — individual nodes and links, like puzzle pieces. Step three: the learner reconstructs the map from those pieces — this is the actual learning activity. Step four: the system compares the learner's map to the goal map and classifies every proposition as matching, missing, or excessive. Because the components are identical, the comparison is exact — no subjective grading. This makes Kit-Build fundamentally different from open-ended concept mapping. It's close-ended, automatic, and gives immediate formative feedback."
Walk the diagram. Use your hand or a pointer to trace top to bottom as you explain each step. This is the only slide where reading the diagram aloud makes sense — it's the architecture of your entire platform.
→ Transition: "Kit-Build has been applied in several contexts — here's where the gap sits."
On the slide: Left: Hirashima (2015), Alkhateeb (2015, 2016), Pailai (2017), Andoko (2020), Funaoi (2011). Right box: "Gap — KB reading studies ≈ English. Japanese adds Kanji, scripts, SOV. Lacking Japanese language learning experiment. No modern web platform + research workflow."
What to say
"The Kit-Build framework was established by Hirashima in 2015. Alkhateeb showed in 2015 and 2016 that Kit-Build produces better delayed recall than summarisation and selective underlining — important because it suggests the benefit is in retention, not immediate scores. Pailai in 2017 developed analytics for formative assessment. Andoko in 2020 applied it to EFL reading and found it outperformed summarisation. Funaoi in 2011 found that retention gains were specific to kit-covered content. But here's the gap: every Kit-Build reading study has been in English. Japanese adds three scripts, Kanji, and SOV structure. There is no published Kit-Build experiment for Japanese language learning. And no modern web platform integrates the full Kit-Build workflow with a research pipeline."
Name-drop with purpose. You're not just listing papers. Each one makes a specific point that builds toward your gap: delayed recall (Alkhateeb), formative assessment (Pailai), EFL reading (Andoko), content-specific retention (Funaoi). Say what each contributed, not just their names.
METHODOLOGY — 4 slides, ~2:45
→ Transition: "With that gap established, here's what I did."
On the slide: One-group pretest-posttest pilot. N students, 2 Marketing Management classes. Beginner Japanese course, convenience sampling. ~X% no JLPT. JLPT self-reported, unverified. Right: stat cards (Enrolled, Analysed, Without JLPT, Mean age).
What to say
"The design is a one-group pretest-posttest pilot. I want to be explicit: no control group. That means this study measures direction, not causation. Participants were beginner Japanese learners from two Marketing Management classes at Polinema — convenience sampling, not random. Thirty-five students enrolled, and after data quality checks, 31 were included in analysis. The vast majority reported no JLPT certification, and where JLPT was reported, it was self-reported and unverified. That's important because it means 'beginner' is defined by course enrolment, not standardised test scores."
Say the numbers on the stat cards. "35 enrolled, 31 analysed" — these are concrete and build credibility. Don't just let them sit there.
→ Transition: "The entire study ran on a platform I built called Yomilink."
On the slide: Modern web-based KB for research. MIT license. Left column: why new platform? (lacks consent, test delivery, export, easy deployment). Right column: KB core (goal map, kit, reconstruction, diagnosis) + research pipeline (consent, pre/post-test, TAM, feedback, analytics, export).
What to say
"Yomilink is a modern web-based Kit-Build platform I developed for this research. It's open-source under the MIT license. The original Kit-Build implementations lacked integrated consent collection, test delivery, data export, and easy web deployment — all of which are necessary for running a study. Yomilink bundles the full Kit-Build core — goal map preparation, kit generation, learner reconstruction, and automatic proposition diagnosis — with a complete research pipeline: consent forms, pre-test and post-test, TAM questionnaire, feedback collection, analytics dashboard, and one-click data export. Everything runs in the browser. No installation required."
Platform as contribution. This is one of your three contributions. Don't rush it. The fact that you built the tool yourself is significant — make sure they know it.
→ Transition: "Here's how the instruments and procedure fit together."
On the slide: Reading test: 10 MCQ, same pre & post, わたしのうち (Minna no Nihongo Ch. 10), furigana + dictionary allowed. Goal map: 7 propositions. TAM: PU + PEoU, 5-point Likert, threshold ≥ 3.5. Feedback: 3 open-ended, thematic analysis. Right: three sessions — Day 1 offline (intro & demo), Day 2 offline (familiarisation + calibration), Day 3 online (registration, pre-test, KB, post-test, TAM, feedback).
What to say
"The reading test used 10 multiple-choice questions based on a passage from Minna no Nihongo Chapter 10 — 'Watashi no Uchi,' or 'My House.' Furigana was provided on all Kanji and a dictionary was allowed throughout. The same 10 questions were used for both pre-test and post-test — I'll address that choice in the discussion. The goal map contained 7 propositions covering spatial and functional relationships in the passage. For technology acceptance, I used Davis's TAM model — perceived usefulness and perceived ease of use — on a 5-point Likert scale with a threshold of 3.5. Three open-ended feedback questions were analysed thematically. The study ran over three sessions: Day 1 was an offline introduction and demo, Day 2 was familiarisation and instrument calibration — the results from Day 2 were excluded — and Day 3 was the main online session where everything happened in sequence through Yomilink."
Mention Day 2 calibration explicitly. It shows you pilot-tested your instruments before the main run. This is a methodological strength, not a weakness.
→ Transition: "For analysis, I used both descriptive and inferential approaches."
On the slide: Descriptive: mean, median, SD, trend counts. RQ1: Shapiro-Wilk, paired t (CLT n≥30), Wilcoxon, Cohen's d, IQR outlier sensitivity. RQ2: Spearman (KB vs Δ), Pearson (pre vs gain), median split → Welch t + Mann-Whitney, Hedges' g. RQ3: construct means vs 3.5, thematic coding. Emphasis: "Direction, not causation."
What to say
"For descriptive statistics, I report mean, median, standard deviation, and trend counts. For the first research question on score direction, I started with Shapiro-Wilk for normality — the data was non-normal, as expected with bounded scores. I used the paired t-test as the primary test, relying on the central limit theorem since n is above 30. Wilcoxon signed-rank served as the non-parametric confirmation. Cohen's d gives the effect size. I also ran IQR-based outlier sensitivity. For the second question on Kit-Build reconstruction, I used Spearman correlation between Kit-Build score and score change, Pearson correlation between pre-test and gain, and a median split with Welch's t-test and Mann-Whitney, with Hedges' g for effect size. For the third question on perceptions, I compared TAM construct means against the 3.5 threshold and used thematic coding for open feedback. Throughout, the emphasis is: direction, not causation."
Don't get bogged down. This slide lists many tests. You don't need to justify each one here — that's for Q&A. Your job is to show you have a systematic analysis plan. Move briskly.
RESULTS — 11 slides, ~5:30
→ Transition: "Now let me walk through what I found, organised by research question."
On the slide: Two classes, beginner Japanese. Mean age, mean study months, prior score mean. ~X% no JLPT. Tests < 60s flagged rushed → N_valid of N analysed. Right: stat cards (Enrolled, Analysed, Excluded, JLPT none).
What to say
"Before the main results, a note on data quality. Thirty-five students participated. I excluded participants whose test completion time was under 60 seconds — suggesting rushed or inattentive responses. This left 31 valid participants for analysis. The mean age was about 20, with roughly a year and a half of Japanese study on average. Prior course scores averaged in the mid-70s out of 100. The vast majority had no JLPT certification. This confirms the sample as genuine beginners with some classroom experience but no standardised proficiency."
Lead with the exclusion. Mentioning data quality first shows rigour. It also preempts questions about "did you check for careless responses?"
→ Transition: "Research question one: what happened to reading comprehension scores?"
On the slide: Box plot: Pre vs Post score distribution. Table: Mean, Median, SD for Pre, Post, Delta. Caption: "Mean 7.06 → 7.77 / 10. High baseline → ceiling effect."
What to say
"The pre-test mean was 7.06 out of 10. The post-test mean was 7.77 — a gain of 0.71 points. The median stayed at 8. The box plot shows the distributions overlapping substantially, which is expected for a small within-group change. Notice the pre-test distribution is already quite high — the median is 8 out of 10. This is the ceiling effect I'll return to. The standard deviations are large — around 2.4 for both pre and post — which means there was substantial individual variation. Some learners already scored near the top; others had more room to move."
Don't just read the table. Point to the box plot and say "notice how the pre-test is already high." Visuals need narration. The audience sees the chart — you tell them what to notice.
→ Transition: "The aggregate hides individual variation — let's look at that."
On the slide: Dumbbell chart: each participant pre → post. Stat cards: Improved (n, %), Same (n, %), Declined (n, %). Avg gain: X pts.
What to say
"The dumbbell chart shows every participant's individual trajectory. Each line is one person — the left dot is their pre-test, the right dot is their post-test. Green lines go up, red lines go down. About 65% of participants improved — that's roughly two-thirds. About 19% stayed the same. About 16% declined. The average gain among those who improved was modest. The key observation: the people who improved the most started lower. The people who declined or stayed flat started higher. This pattern — low scorers gain, high scorers don't — is the signature of a ceiling effect."
Pattern, not just numbers. Don't recite every stat card. Point at the dumbbell chart and describe the shape — upward lines cluster at the bottom, flat lines at the top. That's the story.
→ Transition: "Research question two: how did learners perform on the Kit-Build reconstruction itself?"
On the slide: Histogram of KB scores. Mean ~X%, Median ~X%, Range X%–X%. ~X propositions of 7 reconstructed. Right box: "KB score: structure + extraction ability → broader construct than MCQ."
What to say
"The Kit-Build reconstruction scores averaged around 54%. On a goal map of 7 propositions, that's about 3.8 propositions correctly reconstructed on average. The median was similar, but the range is striking — from zero to one hundred percent. This wide spread tells us that some learners fully grasped the passage structure and could reconstruct it independently, while others struggled significantly. The highlighted point in the box is important: the Kit-Build score captures more than just structural understanding. Because the kit components contain Kanji that learners must decode, the score also reflects extraction ability. This makes it a broader construct than the multiple-choice test, which provides recognition cues. I'll return to this in the discussion."
→ Transition: "Before the inferential results, a note on engagement during the session."
On the slide: Grouped bar chart: Engaged, Moderate, Marginal, Rushed — Pre vs Post. Text: "Engaged X% → Y%. Likely fatigue. 6 phases (~30–40 min) on top of Kanji + novel task. Extraneous cognitive load, not disinterest."
What to say
"Engagement dropped from pre-test to post-test — engaged participants fell while marginal and rushed categories rose. This is almost certainly fatigue. The online session had six phases stacked back-to-back — registration, consent, pre-test, Kit-Build activity, post-test, TAM, feedback — all on top of processing Kanji with a novel task. This is extraneous cognitive load. It's not that learners lost interest in the platform. It's that 30 to 40 minutes of sustained Kanji processing is exhausting for beginners. Future studies should split this across multiple sessions."
Frame it as design insight, not failure. "Likely fatigue" + "extraneous cognitive load" + "split across sessions next time." This transforms a weakness into a methodological recommendation.
→ Transition: "Research question three: what did learners think of Yomilink?"
On the slide: Bar chart of per-item TAM means, dashed line at 3.5 threshold. Stat cards: PU mean, PEoU mean. Caption: "Both above 3.5 → positive initial acceptance."
What to say
"All 35 participants completed the TAM questionnaire. Both perceived usefulness and perceived ease of use had means above the 3.5 threshold — indicating positive initial acceptance. The per-item bar chart shows most items clustered above the threshold line. This is early-stage evidence — it doesn't prove long-term adoption — but it tells us that beginners found Yomilink usable and saw value in it on first encounter. For a pilot study testing a novel platform with a novel task, clearing the acceptance threshold is a meaningful result."
Hedge appropriately. "Initial acceptance" not "adoption." "Early-stage evidence" not "proven." TAM at one time point is a snapshot, not a trajectory.
→ Transition: "Open-ended feedback added qualitative colour to these numbers."
On the slide: Three columns — Positive (helps learning X, simple/easy X), Difficulties (initial confusion X, Kanji difficulty X), Suggestions (more vocab X, better UI X). Caption: "Top positive 'helps learning'; top difficulty 'initial confusion'; Kanji difficulty despite furigana."
What to say
"The most common positive theme — mentioned 24 times — was that the platform helped learning. Learners felt the reconstruction activity made them think about how ideas connect. The most common difficulty was initial confusion — 12 participants mentioned the task was unfamiliar at first. This is expected for a novel activity and reinforces the need for familiarisation sessions, which we included as Day 2. Notably, despite furigana being provided, some learners still reported Kanji difficulty. This suggests that furigana alone doesn't eliminate the decoding barrier — it helps with pronunciation but not necessarily with mapping characters to meaning under time pressure. Top suggestions were more vocabulary practice and better user interface — both actionable for future development."
Don't read every number. Highlight the top theme in each column and the Kanji insight. The rest is on the slide if examiners want to scan.
→ Transition: "Now the inferential statistics — was the score change statistically meaningful?"
On the slide: Shapiro-Wilk all non-normal (expected). Paired t primary (CLT robust at n≥30). Table: Paired t (t = 2.219, df = 30, p = 0.034, d = 0.399), Wilcoxon (W = 52, p = 0.014). Caption: "Significant positive direction, small-to-medium effect."
What to say
"Shapiro-Wilk confirmed that the data is non-normal — as expected with bounded scores and a ceiling effect. The paired t-test, which is robust at n above 30 via the central limit theorem, showed a statistically significant difference: t of 2.219 with 30 degrees of freedom, p equals 0.034. Cohen's d was 0.399 — a small-to-medium effect. The Wilcoxon signed-rank test, which doesn't assume normality, confirmed this with p equals 0.014. I also ran a sensitivity analysis excluding one outlier — a participant who dropped 6 points. With that participant excluded, the t-value rose to 3.52 with p less than 0.001 and a medium effect size. The result holds: there is a statistically significant positive score direction."
Mention the outlier analysis. It shows you didn't just run one test and call it done. Sensitivity analysis is a marker of thoroughness. Examiners notice this.
→ Transition: "Did better Kit-Build reconstruction predict larger score gains?"
On the slide: Scatter plot: KB score vs Delta. Spearman ρ = 0.198, p = 0.284. Caption: "No relation. Japanese artifact: MCQ scaffolds Kanji, KB demands independent decoding. Score = structure + decoding."
What to say
"The Spearman correlation between Kit-Build score and score change was 0.198 — weak and not statistically significant, with p equals 0.284. The scatter plot confirms this — there's no visible pattern. This might seem like a negative result, but it's actually informative. In English Kit-Build studies, reconstruction score and comprehension test score tend to correlate because both measure the same underlying construct: understanding of relationships. In Japanese, the tasks diverge. The multiple-choice test provides answer options that scaffold Kanji recognition — it's a recognition task. The Kit-Build reconstruction requires independent proposition extraction — it's a production task. The KB score captures extraction ability that the MCQ does not. So the weak correlation is not a failure of Kit-Build — it's evidence that orthographic complexity introduces a construct distinction that doesn't exist in alphabetic-language studies."
This is your most sophisticated argument. Practice this paragraph until it flows. It shows you understand not just your results but the measurement construct — a level of thinking examiners respect.
→ Transition: "That ceiling effect I've been hinting at — let's quantify it."
On the slide: Pearson r = −0.549, p = 0.001 (pre vs gain). Median split table: Low pre (n = 19, mean Δ = +1.53), High pre (n = 12, mean Δ = −0.33). Welch t = 2.701, p = 0.013. Hedges' g = 0.988 (large). Caption: "Ceiling confirmed: X% scored ≥ 8 at pre, X% scored perfect 10. Low-pre had room to improve; high-pre decline → test-retest noise."
What to say
"The Pearson correlation between pre-test score and gain was negative 0.549 — highly significant at p equals 0.001. This means the higher someone scored at pre-test, the less they gained. The median split at 8 makes this concrete. Nineteen learners were at or below the median — they gained an average of 1.53 points. Twelve learners were above the median — they actually declined slightly, by 0.33 points on average. The difference between these groups is large: Welch's t of 2.7, p equals 0.013, and Hedges' g of 0.988 — that's a large effect. A substantial proportion of participants scored 8 or above at pre-test, and some scored a perfect 10 — leaving literally zero room to improve. The ceiling effect is real and substantial. It doesn't invalidate the positive direction — if anything, it means the true effect might be larger if measured with more difficult material — but it constrains what we can conclude from this specific test."
Frame the ceiling as a measurement issue, not a study failure. "The true effect might be larger if measured with more difficult material." This turns a limitation into a future research direction.
→ Transition: "Before moving to discussion, here's everything at a glance."
On the slide: Seven bullet points covering: feasible platform, positive direction, KB reconstruction, no KB–gain correlation, ceiling effect, engagement drop, positive acceptance.
What to say
"To summarise the results. One: Yomilink proved feasible — the full Kit-Build plus research workflow ran online with 35 participants. Two: scores moved in a positive direction, from 7.06 to 7.77 out of 10, statistically significant with a small-to-medium effect. Three: Kit-Build reconstruction averaged around 54%, with wide variation. Four: there was no correlation between reconstruction quality and score gain — I've argued this reflects construct contamination from Kanji, not framework failure. Five: a substantial ceiling effect was confirmed — low-scorers gained, high-scorers did not. Six: engagement dropped, likely due to cognitive load from the single-session design. Seven: technology acceptance was positive, with both PU and PEoU above the threshold."
This is your map. If you get lost or rushed, this slide is your safety net. You can read the seven bullets and skip ahead. Don't skip it unless you're severely over time.
DISCUSSION — 4 slides, ~2:45
→ Transition: "Now let me interpret what these results mean — and what they don't mean."
On the slide: Left: Cannot conclude (KB causes improvement, Yomilink outperforms original, results generalise). Threats acknowledged (no control, same items, assisted, familiar passage + ceiling). Right box: Can conclude (Yomilink feasible, score direction positive, pilot procedure tested, positive initial TAM). Caption: "This is a feasibility pilot, every limitation above is why."
What to say
"This is the most important slide for examiners. What can this study conclude? Yomilink is feasible for online Kit-Build with Japanese readers. The score direction is positive, providing evidence to justify future controlled testing. The pilot procedure is tested and documented. Initial acceptance is positive. What can this study not conclude? Kit-Build causes improvement — no control group. Yomilink outperforms the original Kit-Build — no direct comparison. Results generalise beyond this sample — they don't. Every threat to validity is acknowledged: no control group, same pre and post items, assisted conditions, familiar passage, ceiling effect. These aren't accidents — they're features of a pilot design. The purpose of a pilot is to test feasibility and gather initial data for future controlled studies. This pilot achieved both."
Own the "cannot" column. Saying what you can't conclude is a sign of research maturity. Examiners will test whether you know your limits. This slide shows you do.
→ Transition: "Let me unpack the score direction in more detail — there are important caveats."
On the slide: Both tests agree: positive direction after KB. Caveats: no control, same items (but high-pre declined → argues against pure practice), assisted, Day 2 carryover. Schema theory: passage was course material, many had schemas. Right box: "Proficiency threshold: beginners over-rely on pop-up lookups, trapped in bottom-up processing. Furigana + dictionary may be insufficient scaffolding."
What to say
"Both the paired t-test and Wilcoxon agree — the score direction is positive. But there are important caveats. Without a control group, we can't attribute this to Kit-Build. The same items were used twice — but interestingly, the high-scorers declined, which argues against a pure practice effect. If practice alone drove scores up, everyone should have improved. They didn't. The reading was assisted — furigana and dictionary were always available — so this is assisted comprehension, not unaided. And the Day 2 familiarisation session used the same passage, which may have inflated the pre-test baseline. There's also a schema theory angle: the passage was actual course material, confirmed by the lecturer. Many learners already had adequate background knowledge — they understood the house layout before they read about it. This makes the test easier and contributes to the ceiling. The right-hand box raises a deeper point about proficiency: beginners may be trapped in bottom-up processing, over-relying on pop-up lookups. Furigana and a dictionary help with individual words, but they may not be enough scaffolding to bridge from word-level decoding to relational understanding. This is the bottleneck."
The "practice effect counter-argument" is strong. If same questions inflated scores via practice alone, high-scorers should also have improved — they declined. Use this argument. It's one of your best defences against the "same items" criticism.
→ Transition: "The missing correlation between KB score and gain deserves a deeper look."
On the slide: In English KB: reconstruction and MCQ measure same construct. In Japanese: MCQ scaffolds extraction (recognition), KB demands independent extraction (production). KB score captures extraction ability MCQ doesn't → broader construct → weak correlation. Consistent with Alkhateeb (no immediate gain, better delayed recall) and Funaoi (retention on kit-covered content). Feedback aligns: Kanji difficulty + auto-translation requests.
What to say
"The lack of correlation between Kit-Build score and test gain is one of the most interesting findings. In English Kit-Build studies, reconstruction and multiple-choice tests measure the same thing — do you see how concepts connect? Both are about structural understanding. In Japanese, they diverge. The MCQ provides answer options in Japanese — these scaffold Kanji recognition. It's a recognition task. The Kit-Build reconstruction requires learners to independently extract propositions from Kanji-laden components — it's a production task. The KB score captures extraction ability that the MCQ does not. This means the two instruments measure overlapping but distinct constructs, and the weak correlation follows. This interpretation is consistent with prior Kit-Build work. Alkhateeb found no immediate gain from Kit-Build but better delayed recall. Funaoi found retention gains specific to kit-covered content. Both suggest the payoff may be in long-term retention, which this study did not measure. The qualitative feedback aligns — learners reported Kanji difficulty and requested auto-translation features, confirming that extraction was the bottleneck."
This is your deepest analytical contribution. The "construct contamination" argument — that orthography creates a measurement distinction absent in alphabetic studies — is novel. Practice explaining it until you can do it without the slide.
→ Transition: "Finally, let me address calibration, engagement, and the full limitations picture."
On the slide: Calibration: 3 ambiguous distractors + 2 uncovered vocab items revised to 10 course-appropriate items. Engagement drop = cognitive load (6 phases + Kanji + novel task). Fix: split across days. Right: 7 limitations listed (no control, same items, assisted, Day 2 carryover, ceiling, small sample, TAM recency bias, single passage, no delayed test).
What to say
"A methodological strength worth highlighting: the Day 2 calibration session identified three ambiguous distractors and two items with uncovered vocabulary. These were revised into the final 10-item instrument. This is exactly why pilots exist — to refine instruments before the main data collection. For engagement, the drop from pre to post is interpretable as cognitive load — six phases stacked in one session, on top of Kanji processing and a novel task. Splitting the online session across multiple days is the most direct fix. The full limitations list includes: no control group, identical pre and post items, assisted reading conditions, Day 2 carryover inflating baseline, ceiling effect from the familiar passage, a small localised sample, TAM measured at one time point with recency bias, a single reading passage, and no delayed post-test to measure retention. Each of these is addressed in the thesis and each points to a specific improvement for future work."
Don't linger on every limitation. List them concisely and move on. Examiners can ask follow-ups if they want detail. Dwelling too long on limitations makes you sound defensive.
CONCLUSION — 2 slides, ~1:30
→ Transition: "Let me bring this together."
On the slide: Five bullet points: Yomilink feasible, positive score direction, ceiling quantified, no KB–gain correlation (construct divergence), calibration + positive TAM. Bottom box: "Three contributions: (1) modern open-source KB platform, (2) early evidence KB applies to Japanese reading, (3) pilot-tested procedure for future controlled studies."
What to say
"To conclude. Yomilink is a modern, open-source web-based Kit-Build platform — demonstrated feasible for Japanese reading comprehension with 35 learners. The score data pointed in a positive direction — 7.06 to 7.77 out of 10, statistically significant, but cautiously interpreted as a single-group pilot with assisted conditions and a ceiling effect. The ceiling effect was quantified: the median split confirmed that low-scorers gained substantially more than high-scorers, with a large Hedges' g. The missing correlation between Kit-Build score and gain is not a framework failure — it reflects construct divergence driven by Japanese orthography and was underpowered. The calibration session refined the instrument, and initial technology acceptance was positive. This study makes three contributions. One: Yomilink itself — a modern, open-source platform. Two: early evidence that Kit-Build can be applied to Japanese reading, with documented caveats. Three: a pilot-tested procedure that future controlled studies can build on."
End on contributions. The last thing before recommendations should be what you contributed — not what you didn't do. Three concrete items, clearly stated.
→ Transition: "Based on these findings, here's what I recommend."
On the slide: Left column: Future research — larger diverse sample, parallel test forms + delayed post-test, multiple passages. Right column: Practice & development — KB as supplementary formative activity, clarify assisted reading goal, familiarisation time before assessment, clearer tutorial, better app interaction, richer analytics, direct platform comparison.
What to say
"For future research, I recommend a larger and more diverse sample across proficiency levels and study programs. Parallel test forms should replace identical pre and post items. A delayed post-test is essential — both Alkhateeb and Funaoi's work suggests the Kit-Build payoff may be in retention, not immediate scores. Multiple passages would control for passage-specific effects. For practice, Kit-Build makes sense as a supplementary formative activity — especially in non-language programs where learners need to read Japanese but aren't linguistics majors. The task goal should be clarified: assisted reading with dictionary and text access is a valid construct, but it should be named as such. Learners need familiarisation time before assessment — the confusion reported in feedback reinforces this. And the platform itself should receive clearer tutorials, better interaction design, and richer analytics. A direct comparison between Yomilink and the original Kit-Build platform would be needed to make any claim of improvement."
CLOSING — 1 slide, ~15 seconds
On the slide: "Thank You — Questions & Discussion." Your name, NIM, GitHub link.
What to say
"Thank you. I'm happy to take your questions."
That's it. Don't add anything. Don't summarise again. Don't nervously fill the silence. Say thank you, stop talking, and wait. Silence after your presentation feels eternal but it's actually 3–5 seconds. Let the examiners speak first.
Transition Phrases Cheat Sheet
| Moving from… | Say… |
| Cover → Problem | "Let me start with why this matters." |
| Problem → Kit-Build | "So how do we help learners see the connections?" |
| Kit-Build → RQs | "Kit-Build has been studied, but there's a gap — which brings us to my research questions." |
| RQs → Framework | "Let me walk through the Kit-Build framework in more detail." |
| Framework → Related work | "Kit-Build has been applied in several contexts — here's where the gap sits." |
| Related work → Method | "With that gap established, here's what I did." |
| Design → Platform | "The entire study ran on a platform I built called Yomilink." |
| Platform → Procedure | "Here's how the instruments and procedure fit together." |
| Procedure → Analysis | "For analysis, I used both descriptive and inferential approaches." |
| Analysis → Results | "Now let me walk through what I found." |
| Between RQs | "Research question two…" / "Research question three…" |
| Results → Discussion | "Now let me interpret what these results mean — and what they don't mean." |
| Discussion → Conclusion | "Let me bring this together." |
| Conclusion → Close | "Thank you. I'm happy to take your questions." |
Final Reminders
Pacing: If you fall behind, cut from the results section. The methodology and discussion are more important than showing every chart. Slides 15 (engagement), 17 (qual feedback), and 19 (scatter plot) can be reduced to one sentence each if needed.
Don't memorise this script word-for-word. Memorise the structure. Each slide has one core message. Know that message. The exact words will come naturally if you've practiced the structure 3+ times out loud.
Print this document. Not to read during the defence — to review 30 minutes before. The act of scanning the slide headers and bold phrases will activate the neural pathways you built during practice.
Primary Source
Your own thesis paper — re-read the Discussion chapter (Chapter 5) the night before. The language in this script mirrors the language in your paper deliberately. Consistency between your oral presentation and your written thesis signals coherence.
Reminder
If any slide's talking points don't feel natural, tell me and I'll rewrite them in your voice. If you want me to add Q&A prep for specific examiner questions, just ask. This is your defence — the script should sound like you.