When people talk about "research" in cyber security, there is a tendency to reach for a fairly narrow set of examples: vulnerability research, exploit development, bug bounties, conference talks and CVEs.
Those are research activities (or outputs arising from research), but they represent only a small part of what research can mean inside a company whose product is ultimately assurance.
A useful place to start is the Technology Readiness Level, or TRL, scale.
TRLs describe the maturity of a technology from basic principles at TRL 1, through proof-of-concept, validation and demonstration, to an operationally proven technology at TRL 9. The scale originated in technology development and is now used extensively by UK Government and UKRI.
Very approximately:
| TRL | Question |
|---|---|
| 1–2 | Is there an idea here? |
| 3 | Can we demonstrate that it might work? |
| 4–5 | Can we make it work under increasingly representative conditions? |
| 6–7 | Can we demonstrate something approaching a real capability? |
| 8–9 | Can it become, and then remain, an operational capability? |
Somewhere in the middle sits what is often called the valley of death: the gap between demonstrating that something is interesting or technically possible and turning it into something somebody can actually deploy, procure or use.
The UK Government's Digital, Data and Technology Playbook explicitly associates this problem with moving from prototype towards an end product: the testing, iteration and scaling needed to convert technical promise into something usable.
For a consultancy or assurance organisation, however, the valley looks slightly different.
We are rarely trying to manufacture a new physical product.
The equivalent problem is getting from:
"We have demonstrated that we can do something interesting."
to:
"We have a repeatable, supportable and commercially viable capability that improves the assurance we can provide to customers."
That requires considerably more than research producing a clever prototype.
Assurance as the Product
I find it more useful to think about an assurance consultancy (the kind of company I work for) as fundamentally selling assurance.
Assurance here means building justified confidence in some property of a product, service, system or organisation through claims, arguments and evidence.
Some assurance activities principally collect evidence supporting an argument.
A configuration review might establish evidence that a security control has been implemented correctly. A source-code review might provide evidence about the implementation of particular security properties.
Other activities deliberately look for evidence that could defeat those arguments.
Penetration testing and red teaming are obvious examples. Rather than asking only "what evidence supports this claim?", we deliberately ask what evidence might show that the claim, the argument or the assumptions underneath it are wrong.
From that perspective, research within an assurance company should ultimately improve our ability to do one or both of those things.
It should help us ask better questions, construct better arguments, collect better evidence or discover better defeaters.
That is a much wider remit than finding vulnerabilities.
Research Is Shorthand for R&d
When I say research in this piece, I mean research and development: capital R, small d.
That is not a statement about which matters more. It is a description of where the uncertainty sits. In an assurance business the research is the part where the route is unknown, and the development is what carries the result across the valley. From here on I will write Research (R&d) when I mean the function, and research when I mean a specific activity.
Where an organisation has a Research (R&d) function that is separate from business as usual, one thing has to be understood by everybody, the Research (R&d) function included: it exists to support BAU.
That sounds obvious. Its consequences are not.
The alternative is for delivery teams to take these questions on themselves, and there are good reasons that looks attractive. They own the problems. They feel them first. They know which ones actually hurt.
It comes at a cost, though.
A team investigating its own problems sees its own problems. It loses the perspective across a wider set that a function looking at the whole business can hold.
It draws only on its own people. A visible Research (R&d) function can pull in the potentially interested from anywhere in the organisation, which is frequently where the useful idea was sitting.
And, arguably most importantly, internal visibility disappears. Nobody else knows the work is happening, so nobody else can contribute to it, challenge it or reuse it.
That last one has a familiar shape. A delivery team working its own research problem is the person who cannot take their leave: the context, the judgement and the half-built tooling accumulate in one place, and the thing that makes the work valuable is the thing that makes it a dependency. A visible Research (R&d) function is the organisational version of real cover.
The support runs the other way too.
A Research (R&d) function can only work on the problems it knows about, and the problems live in delivery. So the question for anyone on the BAU side is not whether the research function is doing something useful. It is: how are you identifying your problems, and how are you communicating them?
In many ways that should be part of the job.
Every engagement that ended with "we couldn't evidence that". Every test that ran three days over because the tooling did not exist. Every report caveat that says the same thing the last five reports said. Those are research questions in their raw form, and if nobody writes them down and passes them on, they stay as anecdotes told in the pub after the job.
The culture I would want is one where a consultant closing an engagement asks two questions as routinely as they fill in the timesheet: what did we assume that turned out to be wrong, and what would have made this easier? The answers do not need to be well formed. Forming them is what the research function is for.
Research done behind closed doors, with a stated objective of improving things for people, carries two risks. The first is that it misses the objective: what looked valuable in the room turns out to have limited value once it is opened up. The second is that it is never peer reviewed at all.
I personally know the struggle of sharing ideas with people only for the feedback to be rough. That is fine. That is what makes the effort worthwhile.
I don't want something sitting on a shelf that isn't real, and until it has been shared and argued over, it isn't.
What Counts as Research?
The UK's definition of R&D for tax purposes provides an interesting, although deliberately narrower, lens.
Under the DSIT/HMRC guidance, R&D occurs where a project seeks an advance in science or technology and attempts to resolve scientific or technological uncertainty. Crucially, the answer must not simply be readily deducible by a competent professional in the field. The guidance also explicitly recognises system uncertainty: cases where the individual components may be understood, but how they can be combined to achieve the required result is not.
It also recognises something important about experimental development: failure does not mean that research did not take place. A project can still constitute R&D where the advance sought was not achieved.
I don't think the tax definition should become the definition of corporate Research (R&d). Routine analysis, adaptation and optimisation are deliberately excluded from the statutory definition, yet some of those activities can be extremely valuable to an assurance business.
But the concept of uncertainty is useful.
Research is often the work we undertake where the correct route from problem to outcome is not yet known.
Sometimes that uncertainty is scientific or technological.
Sometimes it concerns methodology.
Sometimes the components already exist, but nobody yet knows how to assemble them into a reliable service.
Sometimes the unanswered question is whether the result is economically useful at all.
Those aren't all R&D for tax purposes. They are, however, all questions that a corporate Research (R&d) function may need to investigate.
Research, Innovation and the EngD
I should declare an interest here.
My doctorate is an EngD rather than a PhD: a Doctorate of Engineering from the University of Warwick, in research and development for embedded systems. The EngD was created in the early 1990s because the UK research councils wanted a doctorate that produced industrially useful research engineers rather than academics. You spend most of the four years embedded with an industrial sponsor, working on a problem the sponsor actually has, with a taught programme alongside.
Looked at through the TRL lens, an EngD is a four-year residency in the valley of death. You start from something the sponsor knows is a problem and finish with something they could, in principle, deploy.
The part that shaped me most was not the research. It was the examination.
An EngD thesis has to satisfy two tests at once. It has to make a contribution to knowledge, the same as any doctorate. And it has to demonstrate that the work was of value to the sponsor: that somebody could take it and use it.
Those are different thresholds.
Plenty of work clears one and not the other.
You can produce something new that the sponsor cannot use. You can produce something the sponsor finds enormously useful that contains nothing new at all. Neither, on its own, is enough.
That is why I care about the distinction between research and innovation, and why I think an assurance business should too.
The definition of innovation I work with is the application of knowledge, new or existing, to a new problem space.
Two things follow from that.
The first is that not all research is innovative in its own right. Finding a vulnerability, characterising a protocol or synthesising fifty documents into a playbook may be excellent research. It becomes innovation when the knowledge is applied somewhere it has not been applied before, and the application creates value.
The second is that the definition sets thresholds, even if it does not set them precisely. It forces the question of what is actually new: the knowledge, the problem space, or the combination. And it forces the question of how much value the application creates, because applying existing knowledge to a marginally different problem is not much of an innovation.
The EngD examination asks exactly those two questions.
So, it turns out, does the valley of death.
The Different Research We Actually Do
The first is the kind that cyber security already recognises well: technology and vulnerability research.
This includes vulnerability discovery, exploit research, protocol analysis, reverse engineering, open-source security research and participation in bug bounty programmes.
It is important work. But there is an oddity if we try to map it directly onto TRLs.
A researcher finding a vulnerability in a widely deployed operating system may be researching something that is unquestionably at TRL 9. That doesn't make the work itself "TRL 9 research".
The TRL of the subject of the research and the readiness of the capability emerging from the research are different things.
A new technique for finding that class of vulnerability might begin as a TRL 2 hypothesis, become a TRL 3 proof-of-concept, become a TRL 5 automated prototype, and eventually become a TRL 8 capability embedded into normal assurance delivery.
That distinction matters.
The second category is therefore experimental capability development.
Perhaps we think an LLM can improve some element of source-code review. Perhaps a new static-analysis technique can identify a particular vulnerability class. Perhaps we want to automate part of firmware triage.
We know approximately what we want to achieve, but not necessarily how to get there.
The first implementation probably won't work.
Several competing approaches may need to be built.
Some will fail.
The value of the research is partly in discovering why.
This is exactly where research and development become difficult to separate cleanly. Development becomes experimental when the development path itself contains the uncertainty we are trying to resolve.
Then there is assurance methodology research.
How should a particular security claim actually be evidenced?
What constitutes proportionate evidence?
Which testing techniques are genuinely capable of providing that evidence?
What assumptions are we making about the system?
What potential defeaters should we actively search for?
These questions may produce no vulnerability, exploit or flashy demonstration at all. They can nevertheless materially improve the quality of the assurance being delivered.
Then there is an even less glamorous category: knowledge synthesis.
Take fifty pieces of regulatory guidance, standards, organisational policy and accumulated delivery experience and determine what they collectively tell us.
Identify the common patterns.
Work out where they conflict.
Turn the result into a repeatable playbook for a common class of engagement.
Build reusable test methodologies.
Capture the knowledge that otherwise exists only in the heads of experienced consultants.
I wrote recently about the waste hierarchy as a model for annual leave: Reduce, Reuse, Recycle, in order of effect, with the bottom tier being the only one you can hand to an individual. Knowledge synthesis is the Reduce tier applied to the organisation's own dependencies. Cross-train, document, rotate. Every playbook that captures what one person knows is a single point of failure removed, and the person it removes is usually the one who could not take their leave.
Much of this probably isn't statutory R&D.
It is still research.
And finally there is operational and commercial research.
A tool might work technically while making absolutely no sense commercially.
Does spending £100,000 on a new platform actually allow us to deliver more effectively?
Does automating a particular task save meaningful consultant time once review, false positives, infrastructure and maintenance are included?
Does an interesting research prototype solve a sufficiently common customer problem to justify turning it into a supported capability?
Those questions sit directly in the valley of death.
Writing a Research Question
Knowing the classes is useful because each one produces a different shape of question, a different kind of uncertainty and, crucially, a different kind of output.
A good research question, in any of the five classes, has the same parts.
It names the uncertainty: what we do not know, and why a competent professional could not simply look it up.
It says which class it belongs to, because that tells you who should work on it and what "done" looks like.
It can be turned into a hypothesis: a statement that evidence could confirm or defeat.
It says what capability should exist once the uncertainty is reduced.
And it names the likely output before anyone starts, because a question with no plausible output is a curiosity rather than a research question.
Very approximately:
| Class | Question shape | Uncertainty | Likely outputs |
|---|---|---|---|
| Technology and vulnerability | Does this technology have a weakness of this kind? Can this class of flaw be found with this technique? | Technological | Advisory or CVE, proof-of-concept, tooling, technique write-up |
| Experimental capability | Can we build something that does X to standard Y, and what does it cost to find out? | Development path | Prototype, evaluation against a baseline, a decision to adopt or abandon, a record of what failed and why |
| Assurance methodology | What evidence would support this claim, and what would defeat it? | Methodological | Test methodology, evidence templates, argument patterns, scoping guidance |
| Knowledge synthesis | What do these sources collectively require, and where do they conflict? | Interpretive | Playbook, mapping, training material, reusable scope |
| Operational and commercial | Does this pay for itself once the hidden costs are included? | Economic | Business case, go or no-go decision, pricing, service definition |
Two things about the table are worth noticing.
The first is that the outputs in the right-hand column get quieter as you go down. A CVE announces itself. A service definition does not. That is the next section.
The second is that a question rarely stays in one row. A vulnerability class found in the first row becomes a technique in the second, a methodology in the third, a playbook in the fourth, and a commercial decision in the fifth. That movement is what crossing the valley looks like, and it is the reason the classes belong to one function rather than five.
It is also why I am hesitant to simply encourage more vulnerability research when the request amounts to "this is a cool product". That is not a research question. It is an interest, and interests are where research starts, but the rubric above has five parts and "cool" satisfies none of them.
That does not mean the opposite, either. I am not arguing for less vulnerability research. I am asking for the first step to be done: can you communicate why this is a target of interest? Who depends on it, what claim about it would matter if it fell, what class of weakness you expect and why, and what the organisation would be able to do afterwards that it cannot do now. Answer those and the question usually writes itself. Fail to answer them and the work may still be fun, but nobody will be able to say afterwards whether it was worth doing.
The Classes People Avoid
A reflection, because I think it is worth saying out loud.
Many technically minded people are drawn to the top of that table and quietly avoid the bottom of it. Vulnerability research feels like research. Methodology, synthesis and the commercial questions feel like admin, or management, or somebody else's job.
The awkward part is that the people avoiding those rows are usually the ones holding the knowledge they need.
The consultant who has delivered forty engagements of the same type knows exactly which evidence convinces a regulator and which does not. They know where the guidance contradicts itself, because they have had to pick a side on a live job. They know which parts of the methodology are theatre. That is the raw material for three of the five rows, and most of it has never been written down.
I don't think that is laziness, and I don't think it is disdain. Writing and reflecting are frightening in a way that a debugger is not. A finding is either reproducible or it isn't. A piece of writing exposes how you think, and rough feedback on how you think lands differently from rough feedback on an exploit. I know that well enough myself to recognise it in other people.
This is where the rubric earns its keep.
Getting familiar with research question framing does something useful for the reluctant: it makes the outputs from one row usable as inputs to another. A vulnerability write-up becomes evidence for a methodology question. Twenty engagement reports become a knowledge synthesis dataset. A methodology becomes the thing whose commercial value the fifth row measures. Once you can see the outputs as interchangeable parts, the writing stops being a confession and becomes assembly.
I should be careful with that word, because it makes the writing sound easy, and for me it never was.
Doing the thinking and solving the problem was always my forte. Assembling that into prose was hard. Showing the process end to end, which is what research actually is, was harder still: what I assumed and why, when the assumption stopped being true, which routes I did not take, and how the argument stayed honest as the evidence came in. The solving is the part people mean when they say they enjoy research. The showing is the part that makes it research.
Large language models have made the assembly of prose seemingly trivial, and that has not made the process less important. It has made it more important and harder to get right, because fluent text is now cheap, and a fluent paragraph with an undocumented assumption inside it is indistinguishable, at a glance, from one built on evidence. The scarce skill is no longer producing the write-up. It is knowing what the write-up rests on.
A concrete example. Take the health of internal knowledge and methodology as a metric: how many engagement types have a current playbook, when each was last exercised, how many depend on one named person. That is a synthesis question about existing material, a methodology question about what "current" should mean, and a commercial question about what the gaps cost. Nobody has to write anything new to start answering it. They have to look at what exists, which is the part people avoid.
Research Outputs Are Not Research
This also creates a distinction that I think cyber security sometimes muddles.
A CVE is not research.
A conference presentation is not research.
A vulnerability disclosure is not research.
A blog post is not research.
They are outputs, dissemination mechanisms and sometimes marketing activities arising from research.
They are important. Research that nobody can discover, reuse or learn from has limited organisational impact.
But measuring research by those outputs risks selecting for research which is easy to talk about rather than research which materially improves the organisation.
A carefully constructed delivery playbook might create considerably more long-term value than another conference presentation, while being almost invisible externally.
It is the same problem I described with the leave post. Preventative work registers as an absence of problems, and an absence of problems does not appear in a pipeline report. A conference talk is the recycling of research: real, worth doing, and the only tier that can be handed to one person with a slide deck. The Reduce and Reuse tiers, the playbook and the methodology that stop the next ten engagements going wrong, sit above it and are measured by nobody.
TRLs as a Loop, Not a Conveyor Belt
There is one final problem with applying TRLs too literally.
They imply a journey from TRL 1 to TRL 9.
Modern technology rarely works like that.
A product reaches operational maturity, encounters the real world, and generates new questions.
A vulnerability discovered in a TRL 9 product may expose a previously unexplored scientific or engineering problem.
That produces a TRL 1 or TRL 2 research question.
That research generates a new mitigation.
The mitigation is prototyped, tested and ultimately incorporated into the next product release.
Operational experience therefore feeds back into fundamental and applied research.
The same should happen with assurance services.
Delivery generates evidence about our methodologies.
Research examines that evidence.
New techniques are developed.
They are trialled in controlled environments, piloted on suitable engagements and eventually become normal delivery.
Delivery then exposes the next set of assumptions, limitations and unanswered questions.
The loop also maps naturally onto a development lifecycle such as the V-model, and placing the ISO/SAE 21434 work products on it (tailored, off-the-shelf and newly developed alike) is a post for another day.
Even recent UK Government guidance makes an important qualification here: TRLs describe technical maturity, not whether something is actually ready to be used or sold. Commercialisation additionally depends upon service-delivery capability, compliance, user demand and other factors.
A Worked Example: Large Language Models
All of the above is fairly abstract, so it is worth testing against the topic that currently dominates every conversation about research in this industry.
A security company looking at large language models will almost always start in the same place.
It will start with LLMs as a tool. Can the model triage findings? Can it review source code? Can it write the first draft of a report? Alongside that sits the vulnerability research the industry already recognises: prompt injection, jailbreaks, model and data extraction, poisoned training sets.
That is where the taxonomy predicts we would start, because those are the two classes of research cyber security is most comfortable with. It is also a crowded space, and much of it is about the model rather than about assurance.
The other framing is LLMs as a subject of assurance.
A customer ships a product with a model inside it. A customer buys a service built on one. A customer's organisation quietly adopts one, department by department, and would like to know what that has done to its security posture.
Each of those is a request for justified confidence in a system whose behaviour is probabilistic, whose training data cannot be inspected, and whose supplier may replace the model underneath it without notice.
That is a new problem space, and by the definition of innovation above, applying what we already know about claims, arguments and evidence to it is innovation whether or not any of the underlying knowledge is new.
Walking the five classes through it:
Technology and vulnerability research is well underway, well funded and mostly public. It is necessary. It is not sufficient.
Experimental capability development is where the tool framing lives. The first LLM-assisted code review will not work as hoped. Several approaches will need building. Some will fail, and the value is partly in knowing precisely how.
Assurance methodology research is the hard part. What claim does "our AI feature is secure" actually make? What evidence could support it when the same input may produce a different output tomorrow? Which defeaters should we search for? Non-determinism attacks the repeatability of evidence, and a model update attacks the validity of everything collected before it.
Knowledge synthesis is unglamorous and urgent. Between the EU AI Act, the NCSC's guidance on secure AI development, ISO/IEC 42001, the OWASP list for LLM applications and every customer's internal AI policy, somebody has to work out what these collectively require and where they contradict each other, then turn that into a playbook a consultant can pick up.
Operational and commercial research is the valley itself. Does an LLM save consultant time once every one of its outputs has been checked? What does a hallucinated finding cost when it reaches a customer? And, more fundamentally, can anything an LLM produces be admitted into an assurance argument as evidence, or is it only ever a claim that still needs evidence of its own?
Notice the TRL oddity again. The products are at TRL 9. Most of the methods for assuring them are somewhere around TRL 2 or 3.
Notice the loop, too. The first engagement on an LLM-enabled product will expose an assumption nobody wrote down: that the model tested is the model deployed, that the guardrails are part of the product rather than the supplier's terms of service, that the training data boundary matches the data-protection boundary. Each of those is the next research question.
Some Research Questions
None of these is finished. They are the shape of the uncertainty, which is where research starts, and each is mapped to the rubric above so that what is missing is visible. Every one would need narrowing before it became a project.
What does a security claim about an LLM-enabled product actually mean, and what evidence can support it?
Uncertainty: methodological. The claim is being made about a system that does not behave the same way twice, and the answer is not readily deducible from existing test methods.
Class: assurance methodology research.
Hypothesis: statistical claims over a defined input distribution, with stated confidence bounds, can be evidenced by repeated testing and are more defensible than binary pass or fail claims. Defeated if a model update shifts the distribution after the evidence was collected.
Capability: a way of stating and evidencing security claims about probabilistic components that survives contact with a customer's regulator.
Likely outputs: a test methodology, a claim template, argument patterns for non-deterministic systems, and scoping guidance for engagements that contain a model.
Does LLM-assisted source-code review increase the rate of true findings per consultant hour, once review of the output is included?
Uncertainty: development path. The tooling exists; nobody has measured the net effect on a real engagement, and the route to a version that helps is not known.
Class: experimental capability development.
Hypothesis: recall improves on well-characterised vulnerability classes and does not improve on logic and authorisation flaws, and the net time saving is considerably smaller than the raw output volume suggests. Confirmed or defeated by controlled comparison on codebases with known defects.
Capability: a supported code-review assistant with a known effect size, or a documented reason not to have one.
Likely outputs: a prototype, an evaluation against a manual baseline, an adopt or abandon decision, and a record of what failed and why.
When an organisation adopts an LLM, which new assumptions enter its existing security argument, and which of those can be evidenced at all?
Uncertainty: system composition. The components (people, suppliers, data boundaries, the model) are each understood, but how they combine is not.
Class: assurance methodology research, with a knowledge synthesis component to establish what the guidance currently asks for.
Hypothesis: most of the new assumptions concern people and suppliers rather than the model itself, and existing organisational assurance techniques cover them once the assumptions are made explicit.
Capability: an organisational assurance engagement that accounts for adopted models without starting from a blank page.
Likely outputs: an assumptions catalogue, a mapping to existing controls and guidance, and a reusable scope.
Can output from an LLM be admitted into an assurance argument as evidence, and at what verification cost?
Uncertainty: methodological and economic at once. What counts as evidence is a methodology question; what it costs to make it count is a commercial one.
Class: assurance methodology research feeding operational and commercial research.
Hypothesis: LLM output is a claim rather than evidence, it becomes evidence only when independently reproduced, and the cost of reproduction sets the ceiling on any saving. If that holds, it reshapes the commercial case for the tool framing entirely.
Capability: a rule for when model output may enter an assurance argument, and a costed view of what it takes to get it there.
Likely outputs: an evidence-admissibility guideline, a business case for the tool framing, and quite possibly a go or no-go decision on some of it.
The Useful Question
So perhaps the useful question for corporate Research (R&d) isn't simply:
"What TRL is this?"
It is:
"What uncertainty are we reducing, what capability should exist when we have reduced it, and how does that capability cross the valley into normal operation?"
For an assurance company, that might be the more useful definition of Research (R&d) altogether.