How to Choose a Research Question (Using the Notes You Already Have)
A research question is a promise. To the field, it promises that your article will address something worth knowing. To yourself, it promises that your data can actually deliver the answer.
Most articles that struggle in peer review broke one of those promises right at the start. They asked a question nobody was waiting for, or a question the data could never answer. And the frustrating part is that this happens before you write a single sentence — which means it can also be fixed before you write a single sentence.
I'm a clinical psychologist and statistician. I finished my PhD on schedule, I've published around twenty papers, and I'll be honest with you: I'm a fairly average researcher. What I have is a process. This post is about the first piece of it — how to choose a research question worth an article, using the notes and data you already have.
You won't need new data or new reading. You'll need to organise what you already know, and to resist, at every step, the temptation to let your data decide what the field should care about.
Here's the shape of it. Three moves:
- Listen to the field. Your reading notes already tell you what the field cares about.
- Take stock of your data. A plain-language inventory of your variables and what they can support.
- Match the two. A simple table that pairs field interest against your data until you find the best question you can credibly answer.
📄 Free: I made a fill-in template that walks you through move 3 directly — the Research-Question Matching Worksheet. You list what the field is asking on one side, your variable inventory on the other, and pair them until a question falls out. Grab it here.
Move 1: Start with the field, not your data
Before you open a single dataset, answer one question: does the field actually care about what you found?
This is not a question you can answer by looking at your results. You answer it by listening to the field first — and if you've been keeping a zettelkasten, you've already done most of the work. Your notes are waiting for you.
Two kinds of notes matter here.
Importance notes
Importance notes capture why something matters to the field. They describe prevalence, novelty, and consequences. A good one sounds like this:
"Anxiety disorders are common in the developmental stage of adolescence (12–18 years of age), with a prevalence rate of 4%–8% (Essau et al., 2018; Vizard et al., 2018). Anxiety disorders during adolescence inhibit the ability to seek autonomy and enter adulthood because they negatively affect social interaction, the development of independent living skills and educational outcomes (Swan et al., 2018). Furthermore, these impairments can continue into adulthood if left untreated (Swan & Kendall, 2016)."
Notice what that note is doing: it's making a case that the topic deserves the reader's attention. Prevalence says it's common. Consequences say it matters. Chronicity says it's urgent.
Background notes
Background notes define a concept and summarise what the field already knows about it:
"The best-established treatment for child and adolescent anxiety disorders is cognitive-behavioral therapy (CBT), which has shown effect in specialized settings (efficacy) and in routine-care settings (effectiveness) in several meta-analyses (Whiteside et al., 2020; Wergeland et al., 2020; James et al., 2020)..."
Background notes aren't opinions. They're a summary of where the field currently stands. If you tag your notes by type — importance, background — in Obsidian or Notion, this retrieval step takes minutes.
The two mistakes
There are two ways people get this stage wrong.
The first is peeking at your data too early. You run an analysis, find something interesting, and then go hunting for a reason the field should care. You'll always find an article — but one article is not the same as the field being broadly interested. The result is predictable: reviewers are hard to find and lukewarm when they arrive, and the whole paper feels like it's trying to convince you of something nobody asked about.
The second is the opposite: believing you must read the entire literature first. That's procrastination dressed as rigour. You don't need to know the field in its entirety. You need a general sense — enough to know your topic connects to something the field is actively thinking about. A good sanity check is to ask a colleague or supervisor, would you find this interesting? If they hesitate, listen to that.
I learned this the hard way. Early in my research I was working in a clinical setting and became genuinely excited about multifamily group interventions for anxiety. I found the approach clinically compelling, so when I started writing, I led with it — building the whole argument around why multifamily groups were important and novel. But the field hadn't been asking that question. No matter how interesting I found it personally, I was banging my head against a wall, because I hadn't listened first. The field wasn't waiting for my answer. It had been busy asking a different question entirely.
Go to your zettelkasten before you go to your data. The field has already told you what it finds interesting. Read those notes first.
Move 2: Know your pantry before you cook
Most PhD students treat their data like a lucky dip — they reach in, pull something out, and hope it becomes a paper. The problem isn't that they lack data. It's that they don't know what they have.
Before you can connect the field's interests to your work, you need a clear answer to a simple question: what variables do I have, how were they measured, and under what conditions? That answer tells you what kinds of questions you can actually address — associations, trajectories, causal mechanisms, or something else.
You almost certainly have this information somewhere. The question is whether it's in a form you can use.
Raw data is not an inventory
Raw data files aren't enough. Neither are analysis scripts. What you need is a brief conceptual overview — written in plain language, not in code — that describes your variables at a high level.
Think of it as a pantry inventory. A cook who knows exactly what's on the shelves can plan a meal deliberately. A cook who just opens the fridge and hopes something edible appears is at the mercy of chance. Your data is your pantry. Before you decide what to make, you need to know what's in it.
The format matters far less than the fact that it exists — a table in your lab notebook, a note in Obsidian, a page in Word. What it must contain, for each key variable:
- The name of the variable
- How it was measured or operationalised
- When and from whom it was collected
- What type of question it can support
That last one is the most important. A self-report measure of anxiety at one time point tells you about associations. The same measure at three time points tells you about change. Data from randomised groups tells you about causality. Know which of these you're working with before you start writing.
The two mistakes
Same pattern as before — two ways to trip.
The first is unstructured notes. You know vaguely that you have "anxiety data" and "some parent questionnaires," but not what was measured, in which sample, at which time points. When you sit down to write, the vagueness forces you to dig through files instead of writing, and you can't see what questions are even available to you.
The second is treating raw data as the starting point. Raw data tells you what numbers you have. It doesn't tell you what constructs those numbers represent or what arguments they can support.
Consider John. John reads an interesting paper on emotion regulation and anxiety in adolescents. He checks whether he has anything on that, finds a relevant scale, runs the analysis, and starts writing. Months later the paper comes back from review with a question he wasn't prepared for: why did you not control for family history of anxiety? Your design seems to have collected it. John didn't know he had that variable. It was in the dataset — he'd just never mapped what he had. He'd dipped his hand into the well and come up with one bucket of water, not knowing the well held three.
Map your data before you connect it to the field. You can't find the overlap between what the field wants and what you have if you only know half the equation.
Move 3: The matching table
At this point you have two lists. One is what the field finds interesting, drawn from your importance and background notes. The other is what your data actually contains, drawn from your variable inventory. Your job now is to find where they overlap.
This is a matching exercise, and the sequence matters: you start from field interest and work toward your data, not the other way around. The question is not "what's my best data?" It's "what does the field most want to know, and do I have data good enough to say something about it?"
Build the table
On one side, list the topics and questions your notes suggest the field is actively pursuing. On the other, list your variables and what each can address. Then iterate — looking for cells where high field interest meets acceptable data coverage. You're hunting for the most interesting question the field is asking that your data can credibly answer.
Two things to watch for as you work:
- Look for related clusters. Depression and anxiety, considered separately, may each be moderately interesting. Grouped as internalising disorders, they may address a question the field cares about far more deeply. The table often reveals these groupings when two variables land next to the same field interest.
- Be honest about data quality. Acceptable coverage is not perfect coverage. You want a defensible answer, not an exhaustive one — but "defensible" requires that your data actually measures what you're claiming it measures.
How it fails
There are two ways this exercise goes wrong.
The first is letting the data lead. You pick your strongest variable — the one measured most carefully, with the fewest missing values — and go searching for a field that might care about it. That reverses the priority. You'll find something, but it'll be weakly connected to active field interest, and reviewers will feel it.
The second is claiming a match on superficial similarity. Someone once mentioned they were interested in depression; you have something loosely related to depression; you declare a match. That's not iteration — it's wishful thinking. A real match requires checking that the field interest is genuinely active and that your data genuinely addresses it, not just brushes against it.
I made the second mistake at the start of my own PhD. I'd read a number of articles using symptom severity as the primary outcome — how much anxiety a patient reported before and after treatment. So I wrote my first article the same way, with symptom scores as the main outcome. Reviewers pushed back hard: that wasn't what the field was most interested in. After reading further, I realised that remission — whether a patient lost their diagnosis entirely — was the outcome the field had come to care about most. And I had that data. It had been sitting there the whole time. I'd just never connected what the field was asking with what I actually had, because I let the literature I happened to read first dictate my choices instead of systematically mapping field interest against my data.
Iterate until you find the most interesting question the field is asking that your data can credibly answer. That is your article.
What you have now
You now have something most people don't have when they sit down to write: a research question with a pedigree.
It came from the field's own interests, documented in your notes. It's been checked against an honest inventory of your data. And the connection between the two is explicit — you can state it in one sentence and defend it to a supervisor, a reviewer, or yourself in six months.
Three moves got you here:
- You listened to the field before looking at your data, because the field decides what's interesting — not your dataset.
- You mapped what you have, because you can't match what you can't see.
- You built the matching table, iterating until the strongest field interest met data good enough to address it.
What you don't yet know is whether this question survives contact with other people — and you shouldn't write an entire article to find out. That's a job for the right readers, brought in at the right stage, asked the right way. But the promise is already half-kept: you've chosen a question the field is waiting for, and one your data can answer.
📄 Free: If you want to run move 3 today, download the Research-Question Matching Worksheet — a fill-in template that pairs your field interest against your variable inventory, one row at a time, until a defensible question emerges. Get the worksheet.
Member discussion