QCE Psychology

QCE Psychology IA2 student experiment: how to get top marks

Your IA2 is worth 20% of your final result and is marked out of 20, five marks each across four criteria: Forming, Finding, Analysing, and Interpreting and Evaluating. What moves you from the middle band to the top band is a short list of adjectives in the marking guide, and above all the statistics behind your analysis: a measure of central tendency chosen for your data, a dispersion measure to match, and an inferential test reported properly.

This guide describes the Psychology 2025 v1.0 syllabus, which applies to Year 11 from 2025 and Year 12 from 2026, so the 2026 Year 12 cohort is the first assessed under it. The exemplars quoted below come from the four subject reports published so far (2022 to 2025 cohorts), every one of them marked under the superseded Psychology 2019 v1.3 marking guide. Their criterion names and mark splits describe the old rules. What separates a strong response from a weak one comes down to data and reasoning, so that part survives the revision.

What the task actually asks

You modify an experiment relevant to Unit 3: refine, extend or redirect a practical or simulation you have already performed in class, to address your own related research question or hypothesis. That last part is a syllabus rule. The methodology and research question have to grow out of something done in class, so you cannot arrive with an experiment invented from scratch.

QCAA designs the task around roughly 10 hours of class time. All six assessment objectives are assessed here, against three of the six in the IA1 data test.

Conditions. Up to 2000 words written, or up to 11 minutes multimodal.

What can be done as a group: identifying the experiment, developing the research question, conducting the risk assessment, running the experiment, and collecting data. Processing, analysing, interpreting and writing up are yours alone.

Where each criterion draws from

CriterionAssesses objectivesMarks
Forming1, 2, 65
Finding65
Analysing2, 35
Interpreting and Evaluating4, 55

One thing here catches students out every year, because the criterion names suggest the opposite arrangement. Your methodology is marked in Finding, not in Forming. Forming marks whether you justified your modifications. Finding marks whether the methodology you ended up with actually collects sufficient, relevant data. A well-argued rationale attached to a design that gathers too little usable data scores in two different places, and only one of them rewards the argument.

Forming (5 marks)

Marked on your rationale, your modifications, your research question, and your use of genre, referencing and scientific language.

4-5 marks2-3 marks
a considered rationale for the experimenta reasonable rationale
justified modifications to the methodologyfeasible modifications
a specific and relevant research questiona relevant research question
appropriate use of genre and referencing conventionsbasic genre and referencing conventions
fluent and concise use of scientific language and representationscompetent use of scientific language

Genre conventions, referencing and scientific language all sit inside Forming for IA2. Students working from older material go looking for a communication mark, and there is not one. Do not carry this arrangement across to IA3 either, which distributes those characteristics differently. See the IA3 research investigation guide.

On the rationale, the 2024 report describes the top-band version as one where "a considered rationale clearly connected the research question to Unit 3 subject matter, using key terms and concepts." So name the memory, learning or sensation and perception concepts your experiment tests, and use them.

On modifications, the 2022 report is the most concrete of the four: modifications were justified "with specific reference to concepts from Unit 3 subject matter… by identifying how they would improve the experiment's validity or reliability." Naming what you changed is the easy half. The mark is in saying which of reliability or validity it improves, and how.

Finding (5 marks)

Marked on your methodology, your management of risks and ethical issues, and the raw data you collect.

4-5 marks2-3 marks
a methodology that enables the collection of sufficient, relevant dataa methodology that enables the collection of relevant data
considered management of risks, ethical issues and environmental issuesmanagement of risks, ethical issues and environmental issues
collection of sufficient and relevant raw datacollection of relevant raw data

The whole gap on the middle line is the word considered, and three reports agree on what it means: strategies plus the reasons for them. The 2025 report says considered management "was communicated using explanations of why particular strategies were used." The 2024 report asks for "the discussion of how and why a variety of risks and ethical issues were managed." The 2023 report calls it a discussion that demonstrates "careful and deliberate thought."

In Psychology this criterion has real content, because your participants are people. Consent, the right to withdraw, anonymity and safe stimulus levels are all live issues in a normal Unit 3 experiment, and QCAA's own exemplars below turn on exactly those.

Analysing (5 marks)

Marked on how you process your data and what you find in it. Data processing is the single most repeated piece of IA2 advice across the four reports held, which makes this the criterion to get right first.

4-5 marks2-3 marks
correct and relevant processing of databasic processing of data
thorough identification of relevant trends, patterns and relationshipsidentification of obvious trends
thorough and appropriate identification of the uncertainty and limitations of evidencebasic identification of uncertainty and/or limitations

What correct and relevant processing means

Assembled across the 2022, 2023 and 2024 reports, it means four things, and a response missing any of them is not in the top band:

  1. The most appropriate measure of central tendency, mean or median, chosen for the data you actually have.
  2. A measure of dispersion aligned to it: standard deviation, standard error, or interquartile range.
  3. The most appropriate graphical representation. The 2024 report's example is a bar graph of mean values with error bars showing standard error of the mean or 95% confidence intervals.
  4. A suitable inferential test reported with its p value, testing your null hypothesis. The 2023 report names one- and two-tailed t-tests and the Mann-Whitney U test.

The 2025 report adds a fifth requirement aimed squarely at the new Analysing criterion: the statistical analysis, the methodology and the graphs all have to line up. Its wording is that relevant processing "shows alignment between the correctly performed statistical analysis appropriate for the methodology and the graphical representations used to derive relevant trends."

The strongest single illustration of that distinction in any held report is a 2025 exemplar that plots both. On the first chart: "the SD error bars are elongated, indicating that the data has high variability. This conclusion is supported by the results as the SD for the NPR group is ±15.977 and the SD for the HIPR group is ±18.175, both of which are high numbers that show that the variance from the mean is great. This suggests that there were inconsistencies in the results which impacts the reliability of the experiment." On the second: "Figure 2 shows the uncertainty of the results from the tested conditions through 95% margin of error (MOE) confidence intervals."

Interpreting and Evaluating (5 marks)

Marked on your conclusion, your discussion of reliability and validity, and your improvements and extensions.

4-5 marks2-3 marks
justified conclusion/s linked to the research questionreasonable conclusion/s relevant to the research question
justified discussion of the reliability and validity of the experimental processreasonable description of reliability and validity
improvements and extensions logically derived from the analysis of evidenceimprovements and/or extensions related to the analysis of evidence

Two of those lines turn on a single word each, and both are easy to lose without noticing.

Description against discussion. Naming your reliability and validity issues is a description, and descriptions sit in the middle band. A discussion explains how your design produced those issues and points at the evidence that shows it. The 2023 report adds the naming rule: say which kind you mean. Inter-rater reliability, test-retest reliability, internal validity, external validity. "The experiment was reliable" identifies nothing.

The distinction underneath the two terms is the one QCE science works to throughout. Reliability asks whether another experimenter would obtain the same results. Validity asks whether you measured what you set out to measure, which sends you back to the purpose you stated in your rationale.

Every suggestion has to be derived. All four reports say this, in four different years, and it is the most repeated message in the IA2 commentary after data processing. The 2023 report asks for suggestions "logically derived from uncertainty and limitations identified in the analysis of evidence, e.g. measures of dispersion, confidence intervals, sample size or composition, assumptions of statistical tests." The 2025 report asks for improvements and extensions "considered by referring directly to their effect on related issues with reliability and validity that were identified in the analysis of evidence." An improvement that does not trace back to a limitation you actually found is not top band, no matter how sensible it sounds.

On the conclusion, compare two 2024 excerpts. The strong one justifies its choice of test before reporting the result: "For inferential statistics, a parametric independent t-test was selected to determine the statistical significance of the data (t=1.71, p=0.008). This type of t-test was conducted as the experiment design was independent groups with a directional alternative hypothesis. Since the p-value is less than the alpha level of 0.05, the data can be deemed statistically significant. Therefore, the null hypothesis is rejected, and the directional alternative hypothesis is accepted." The weaker one reports a number and stops: "The p value is 0.48392 which means we will accept the null hypothesis."

What four years of reports and confirmation data show

Data processing is where schools are most generous and QCAA is most strict. Across the 2025 cohort, marked under the superseded 2019 criteria, the one covering data processing was called Analysis of evidence. School judgments on it agreed with QCAA on 86.08% of samples and were reduced on 12.66% of them, the worst agreement of any IA2 criterion in any of the four years held. Overall IA2 agreement that year was 79.32%. Where QCAA changed a mark, it nearly always changed it downward: across all four years the share of marks raised never exceeded 1.82% on any criterion.

QCAA has not published a criterion-by-criterion mapping between the 2019 and 2025 marking guides, so read that finding as a signal about data processing itself, without attaching it to any 2025 criterion. The message it carries is the same either way. If your school has marked your processing highly, that is the part of your response most likely to come down at confirmation, so it is the part to strengthen first.

The Psychology study guide covers how confirmation moves marks across all three internal assessments, and where marks are lost across the subject.

Common mistakes

Putting method limitations in Analysing. Limitations there are about your data. Method critique belongs in Interpreting and Evaluating, where it is the point.

Defaulting to the mean. The criterion asks for the most appropriate measure of central tendency. If your data is skewed or carries an outlier, the mean is the wrong choice, and choosing it costs you on the line about correct and relevant processing.

Reporting a p value without justifying the test. Which test you ran, and why that test suits your design and your hypothesis, is part of correct processing.

Describing reliability and validity instead of discussing them. Two named issues with no explanation of how the design produced them sits in the middle band by definition.

Treating reliability and validity as one idea. They are marked on the same line and they ask different questions.

Offering only an improvement or only an extension. Both are required for the top band, and a strong one of either kind will not cover for the other's absence.

Choosing a practical that yields categorical data. QCAA has named this twice. Without ordinal, interval or ratio measurements you cannot produce the statistics the Analysing criterion is asking for.

Weak versus strong wording

WeakStrongSource
"Participants gave their consent and could leave at any time.""Participants were given a consent form to be signed by a parent/guardian, as they were under 18, to inform them of their child's participation in the experiment and allow them to decide if they wished for their child to participate. Participants were also given the right to withdraw at any time during the experiment, ensuring voluntary participation and control over their involvement."2025 report, p. 16
"Does processing affect memory?""Does the level of processing, structural or semantic, affect the total number of words recalled amongst 15-17 year olds?"2025 report, p. 16
"The mean number of words recalled was calculated.""Although the dependent variable (number of words recalled) was interval-ratio data, a visual inspection of the histograms indicated that the data was not normally distributed… Therefore, the most appropriate measure of central tendency and measure of dispersion were median and interquartile range respectively."2022 report, p. 16
"The p value is 0.48392 which means we will accept the null hypothesis." (2024 report, p. 16)"For inferential statistics, a parametric independent t-test was selected to determine the statistical significance of the data (t=1.71, p=0.008). This type of t-test was conducted as the experiment design was independent groups with a directional alternative hypothesis. Since the p-value is less than the alpha level of 0.05, the data can be deemed statistically significant."2024 report, pp. 13-16
"The error bars were quite large.""the SD error bars are elongated, indicating that the data has high variability… the SD for the NPR group is ±15.977 and the SD for the HIPR group is ±18.175… This suggests that there were inconsistencies in the results which impacts the reliability of the experiment."2025 report, pp. 18-20
"More participants would improve the experiment.""harder, more complex comprehension tests to eliminate the ceiling effect, providing a more sensitive measure for data collection to better discriminate between groups."2025 report, pp. 17-18
"The experiment was reliable.""inter-rater reliability", "test-retest reliability", "internal validity", "external validity", each named and then explained against the evidence2023 report, p. 18

Before you submit

  • Does your rationale connect your research question to named Unit 3 concepts, using their proper terms?
  • Does every modification say whether it improves reliability or validity, and how?
  • Does your research question name both variables and your population in one sentence?
  • Did you choose your measure of central tendency and dispersion for the data you got, and can you say why?
  • Have you run an inferential test, reported its p value, and justified why that test suits your design?
  • Do your error bars show what your caption says they show, variability or uncertainty?
  • Are your stated limitations about the evidence, with method critique saved for Interpreting and Evaluating?
  • Does every ethics and risk strategy come with the reason it is there?
  • Do you name which reliability and which validity you mean, then explain them against your results?
  • Have you given both improvements and extensions, and does each one trace to something in your analysis?

If you are already into Unit 4, the IA3 research investigation asks you to evaluate a claim using other people's evidence instead of your own experiment, which tests a skill IA2 never touches. The Psychology study guide sets out where the marks sit across the whole subject.

Frequently asked questions

How long should the QCE Psychology IA2 be?

Up to 2000 words written, or up to 11 minutes multimodal. QCAA designs the task around roughly 10 hours of class time, drawn from Unit 3. The 2025 syllabus changed these from ranges to ceilings, so 2000 words is a limit and not a target.

What statistics do I need in my Psychology IA2?

Four things, named across the 2022, 2023 and 2024 subject reports: a measure of central tendency chosen for your data (mean or median), a measure of dispersion aligned to it (standard deviation, standard error or interquartile range), a graph that suits the analysis, and an inferential test reported with its p value to test your null hypothesis. Reports name one- and two-tailed t-tests and the Mann-Whitney U test as examples.

Is there still a Communication criterion in QCE Psychology IA2?

No. The 2019 marking guide had Communication as its own 2-mark criterion. In the 2025 IA2 marking guide, genre conventions, referencing and scientific language are marked inside Forming, and all four criteria are worth 5 marks each.

Can I do my QCE Psychology IA2 in a group?

Parts of it. The syllabus allows identifying the experiment, developing the research question, conducting the risk assessment, running the experiment and collecting data to be done as a group. Processing the data, analysing it, interpreting it and writing it up must be your own individual work.

Do standard deviation and standard error bars show the same thing?

No, and the 2025 subject report singles this out as something schools should teach explicitly. Standard deviation bars show variability in your data. Standard error bars and 95% confidence intervals show uncertainty about the mean. Labelling one as the other is a processing error in the Analysing criterion.

Sources

  • Psychology 2025 v1.0 General Senior Syllabus, QCAA, January 2024, pp. 41-44. IA2 specifications, conditions, mark allocation and the instrument-specific marking guide.
  • Psychology 2019 v1.3 General Senior Syllabus, QCAA, June 2018, Section 4.7.2. The superseded IA2 marking guide, used here only to state what the 2025 criteria replaced.
  • Psychology subject report, 2025 cohort, QCAA, January 2026, pp. 13-20. The ethics, research question, improvements and error-bar exemplars, and the variability versus uncertainty warning.
  • Psychology subject report, 2024 cohort, QCAA, January 2025, pp. 12-16. The four-part account of correct and relevant processing, and the justified inferential conclusion.
  • Psychology subject report, 2023 cohort, QCAA, January 2024, pp. 14-19. The naming rule for reliability and validity, the ethics exemplar, and the practical to avoid.
  • Psychology subject report, 2022 cohort, QCAA, February 2023, pp. 13-17. The exemplar choosing median and interquartile range for non-normal data.

Syllabus and assessment material referenced in this guide is © State of Queensland (Queensland Curriculum and Assessment Authority), licensed under CC BY 4.0. See our QCAA licensing notice. AusGrader is an independent study tool and is not affiliated with, endorsed by, or operated by the QCAA.

Keep reading

Practise what you just read

Work through real QCAA Psychology questions and get your written answers marked against the official criteria, instantly.