Lesson 3 Defense Prep

Methodology Defense

Answer questions about design, data cleaning, and procedure

Methodology questions are the most predictable part of a defense. Examiners want to know: did you make defensible choices, and do you understand why?

Research Design Questions

Why did you use a one-group pretest-posttest design instead of a true experiment?
This was a feasibility pilot. The primary goal was to validate Yomilink's workflow, test instruments, and establish score direction before committing resources to a controlled experiment. A single-group design is appropriate for preliminary investigation and aligns with the study's exploratory nature.
Can you claim that Kit-Build improved reading comprehension?
Only cautiously. The study shows a positive score direction after the Kit-Build activity. Because there was no control group, we cannot fully separate the Kit-Build effect from practice effect, material familiarity, or repeated exposure. The safe claim is that students showed improved scores after using Yomilink.
Why didn't you use a control group?
This study was designed as a preliminary pilot and feasibility study. The goal was to validate the platform workflow, instruments, and procedures, and to gather initial evidence on score direction. A controlled experimental design is the natural next step, informed by this pilot.

Data Cleaning Questions

Why is <60 seconds the threshold for invalid tests?
The test consisted of ten reading comprehension questions with a reading passage. Completing the entire test in under 60 seconds means less than six seconds per item, excluding the time needed to read the passage and answer choices. Such responses are implausibly fast for effortful reading comprehension and were flagged as low-engagement responses.
For invalid post-tests, why carry forward the pre-test score instead of excluding the participant?
This is a conservative normalization approach. Carrying forward the pre-test score keeps the participant visible in the dataset while preventing rushed post-test responses from producing artificial score gains. It assumes no observed improvement when the post-test response is invalid, which is more conservative than dropping the case entirely.
Doesn't single-value imputation introduce bias?
Yes, it can. This is acknowledged as a limitation. For a preliminary study, we prioritized keeping participants visible in the data-quality table over optimal missing-data handling. Future work should use cleaner data collection, stronger engagement controls, and sensitivity analyses comparing raw, excluded-case, and normalized datasets.

Procedure Questions

Why was Day 2 (familiarization) conducted with the same material as Day 3?
Day 2 served as platform familiarization and calibration. Using the same material ensured participants understood the task format. The feedback from Day 2 — including difficulty feedback that led to instrument revision — would not have been possible with different material. The 5-day gap between sessions reduced direct recall, though material familiarity carryover is acknowledged as a limitation.
Could Day 2 practice have inflated Day 3 scores?
Possibly. This is acknowledged as a limitation. However, Day 2 was primarily interface familiarization, not comprehension practice. The 5-day gap, combined with the fact that the task was about concept map reconstruction rather than memorization, reduces (but does not eliminate) direct carryover effects.
Why allow dictionaries and passage access during the test?
The study measured assisted reading comprehension — how beginners perform when script decoding is partially scaffolded. This is ecologically valid because real-world reading for beginners typically involves dictionary and reference access. Measuring unaided comprehension of raw Japanese text would test memorization and script decoding, not structural understanding of the text's meaning.
Why use the same questions for pre-test and post-test?
This kept item difficulty constant for a preliminary within-subject comparison. Similar designs appear in related Kit-Build evaluation studies. The limitation is possible practice effect, which is why future research should use parallel test forms. This trade-off was acceptable for a feasibility pilot where controlling item difficulty was prioritized.

Sample Questions

Why only 35 participants? Is that enough?
For a feasibility pilot, N=35 is adequate to establish score direction, platform feasibility, and initial acceptance signals. The study was not powered to detect small effect sizes or test correlational hypotheses with high confidence. The purpose was to generate evidence for a future controlled study with a larger sample.
Can results generalize to all Japanese learners?
No. Participants came from two Marketing Management classes at one institution. The sample represents beginner learners in a specific educational context. Generalization to other populations, proficiency levels, or institutional settings requires further research.
Why describe participants as "beginner Japanese learners" instead of "JLPT N5"?
Because 68.6% of participants had no JLPT certification. Proficiency was inferred from course enrollment and self-report, not formally verified. Describing them as "beginner Japanese learners" is accurate and avoids claiming a proficiency level that was not independently confirmed.
Pattern for all methodology questions:
State what you did → explain why it was defensible for a pilot study → acknowledge the limitation → describe what future work should do instead.

Primary Source

Read Creswell & Creswell (2018), Chapter 10 — pre-experimental designs, threats to validity, and design-specific recommendations.

Reminder

Ask me anything. If you want to practice answering a specific methodology question out loud, just tell me the question and I'll play the examiner.