DevSecOps16 min

Tests that actually catch errors: beyond coverage and a green quality gate

By Dorian Chávez · founder of Hábil and integration architect ·

High coverage and a green quality gate aren't enough: test types, doubles that hide failures, mutation, AI and documentation for tests that catch errors.

Illustrative scenario. A fintech's payments team opens the dashboard on a Friday: test coverage stands at 91%, the code quality analysis shows green and the pipeline —the automated chain that builds, tests and delivers the software— finished without errors. On Monday, a group of customers reports duplicate charges. Nobody lied and no tool failed: each one measured what it knows how to measure. What nobody had measured is whether the tests, of which there were many, would have warned about that particular defect.

This article is for whoever is accountable for that difference: the CTO or technology director of a bank, a fintech, an insurer, a retail chain or a regulated company. It explains what coverage and the quality gate (a set of conditions that the code must meet in order to move forward) really measure, which types of test exist and which failure each one catches, how to check that a test is any good, what role artificial intelligence and code documentation play, and where to start. It includes what we have measured on our own platform, because some of those findings were uncomfortable.

What a green dashboard really measures

It is worth starting with what the tools themselves say.

Coverage answers a single question: was this line of code executed while the tests were running? SonarQube, a code quality analysis tool, defines line coverage as covered lines divided by executable lines, and condition coverage as the true and false branches traversed divided by twice the number of conditions [11]. Executing a line is not the same as checking that it does the right thing: a test with no verification at all (an "assertion," the comparison between what was expected and what happened) can raise coverage without catching anything.

SonarQube's quality gate is a set of configurable conditions. The default one, called "Sonar way," asks for the following on new code: no new issues, coverage of at least 80%, duplication of 3% or less and security hotspots reviewed [10] [Sonar figure: vendor defaults]; each organization adjusts them. Line coverage only says whether a line was executed; Sonar's overall coverage metric combines lines and conditions. In addition, Sonar doesn't generate the coverage: it imports it from another tool that runs the tests before the analysis [12]. From what it measures, what it doesn't do can be deduced: it doesn't execute a transfer, a reconciliation, a policy issuance or an e-commerce payment. That reading is ours, not a sentence from Sonar, but it is consistent with what its documentation describes.

Nor is there a magic coverage figure. Martin Fowler, a leading authority in software engineering, considers it a tool for finding untested code, not a goal; he finds coverage in the high eighties or nineties reasonable, treats 100% as a warning sign and cautions that imposing a minimum invites writing empty tests to reach it [9]. A classic academic study, by Inozemtseva and Holmes (2014), found that coverage does not have a strong correlation with the effectiveness of a test suite once its size is accounted for [21].

And "green" means different things depending on who configures the gate. In an internal review of our platform we found that the analysis of a newly created project passed with 3 active conditions, while that of a mature one demanded 14. The same color, two very different levels of rigor.

The two most common forms of false green

1. Doubles that answer what you expect

To test quickly, teams replace the external pieces —the database, the payment provider, another service— with test doubles: imitations that respond in a predefined way. Fowler distinguishes several types: stubs (canned answers), fakes (simplified but working implementations), spies (which record how they were used) and mocks (which come with preprogrammed expectations) [8]. They are useful for keeping tests fast and isolated.

The risk is simple to state: if the double answers what the programmer believes the real system answers, the test passes even if the real system answers something else. A double can reply "200 OK" while the real provider rejects the header, the encoding, the certificate or the date format. Fowler also points out that tests based on mocks end up more coupled to the implementation: if you reorganize the code without changing what it does, those tests break; and if the real system changes, they keep passing [7].

That is why tests must be taken to real dependencies where the decision is made by the dependency and not by your code: the database with its real engine, the message bus with its own, the contract with the provider verified against the provider. For dependencies you can containerize, tools such as Testcontainers spin up a real, disposable database or queue while the test runs and destroy them when it finishes; its own guide proposes replacing in-memory databases with the real database [13]. For external providers, keep using a sandbox, a contract or an observable double: a container does not replace the credentials, certification or authorization of a real third party.

2. Tests that bless the error

A test written from what the code does, and not from what it should do, certifies the defect. A simple example: the business rule says that amounts of 10,000 pesos or more are blocked, but the code blocks only those greater than 10,000, and the test, written by looking at the code, verifies the current behavior. There is full coverage, everything is green and the rule is violated. It is a known risk of tests written by people and, as we will see, a greater risk in those written by an AI.

Neither of the two forms shows up on the dashboard. They show up when you ask the test suite different questions, and that is what the rest of the article is about.

Test types and the failure each one can catch

No type of test, on its own, covers all the relevant risks. Each type can provide evidence that the others do not provide for that flow. This is the list we use to explain it to a committee:

Test types and the failure each one can catch
TypeWhat it runsWhat it can catchWhat it doesn't seeExample tools
UnitAn isolated function or classCalculations, boundaries, validations, local rulesWhether the pieces fit togetherJUnit, pytest, Go testing, Jest, Vitest
Integration with real dependenciesYour code plus a real, disposable database, bus or serviceBroken migrations, queries that fail against the real schema, serialization, configuration, transactionsComplete user flowsTestcontainers
ContractThe consumer and the provider of an API, each against what was agreedA renamed field, a changed type, a different header that would break the consumerThe provider's internal logicPact
API collectionHTTP requests against an environment, with checksResponse codes, formats, authentication, authorization, regressions of a serviceInternal behaviorPostman, Postman CLI, Newman
Browser (end to end)A simulated person, the browser, the frontend and the backend integratedRoutes, cookies, login, rendering, critical flowsFine details; they are slow and brittle if overusedPlaywright
MutationYour suite against deliberately damaged versions of the codeTests that execute code without detecting changesWhether the business rule itself is the right onePIT (Java), Stryker (JavaScript/TypeScript, C#, Scala)

Some clarifications that matter for deciding:

  • Contract. Pact verifies the provider that the team controls; it does not necessarily validate the real third party. It works "consumer-driven": only what the consumer actually uses is tested; the consumer generates a file with its expectations and the provider verifies it against the real service [14]. It avoids setting up costly integrated tests between two teams. Its own documents state its limits: it is not suited to functional or load tests, to public APIs, or when you don't control both sides [14].
  • API collections. Postman lets you write checks in JavaScript per request, folder or collection [15]. A practical fact that few mention: Newman, the command-line collection runner, remains available for existing flows, but its official repository states that active development is limited to essential maintenance and that, for new flows, they recommend the Postman CLI [16]. If you already have collections running in Newman, there is no urgency; if you are about to start, it makes sense to start with the current tool.
  • Browser. A browser test is not always end to end: it can simulate third parties. Playwright covers Chromium, WebKit and Firefox, and its best practices call for testing what the user sees, isolating each test and simulating third-party sites instead of testing them [17]. It is the type closest to real use, and that is why it is reserved for a few high-value journeys.
  • Mutation. It can be explained in one sentence: it is deliberately sabotaging the system to see whether the alarm sounds. It is developed further below.

Tools by layer

Test types and the failure each one can catch
LayerExample toolsWhat for
Java backendJUnit, Testcontainers, PITBusiness rules, integration with infrastructure, mutation of critical code
Python backendpytest (with its fixtures, reusable supporting data and resources)Rules, parameterization, integration
Go backendtesting and go test, part of the standard library; supports fuzzing tests (random inputs)Units, integration, unexpected inputs
Frontend (JavaScript/TypeScript)Jest or Vitest, with Testing LibraryInterface and component logic, tested the way a person uses them
Frontend, complete journeysPlaywrightCritical flows and cross-browser compatibility
JavaScript/TypeScript, mutationStrykerJS on top of Jest or VitestDetecting weak tests

Testing Library states the principle clearly: the more your tests resemble the way your software is used, the more confidence they can give you [20]. Jest states that Vite is not officially supported and suggests Vitest in that case [20]; it is an example of why a tool is chosen according to the rest of your stack, not by fashion.

Why more types of test pay off more than more tests of the same type

The right claim is not "the more types, the better" without limit. It is this: for a given risk, it makes sense to invest in the type of test that covers a class of failure that nobody has covered yet.

Ten more unit tests on a rounding calculation are not going to uncover a failing database migration, an incompatible HTTP contract, a session that gets lost between screens or a dependency that takes too long. Other types detect those. The research evidence on test diversity points that way: criteria that distinguish different behaviors can find failures that traditional criteria don't see, at an additional cost [23].

The other half of the balance is not going to the opposite extreme. Browser tests are the closest to real use, but also the slowest and most brittle. Fowler describes them as brittle, expensive to write and time-consuming to run [1]. In Google's study of about 4.2 million of its own tests, the largest tests were more flaky, that is, they pass and fail without the code changing; WebDriver and the Android emulator showed above-average rates among the tools analyzed, and the author himself notes that correlation is not causation [5] [Google figure, 2017; not universal]. Its guide recommends, roughly, 70% small tests, 20% integration and 10% end to end [4] [Google figure, 2015; not universal]. It is one company's guidance, not a standard; there are those who argue that discussing percentages is a distraction [3]. Fowler and Google recommend reserving broad tests for the risks that justify them; the proportion depends on the system, and Fowler warns that the pyramid has exceptions [1][4].

An example for a committee: three thousand unit tests with doubles, with no contract tests, mean that a field renamed in an API breaks all its consumers in production without any dashboard warning anyone. Those same three thousand, without real integration, mean that nobody tested the query against the real schema. And if all of the above exists but no test walks through the complete flow, nobody knows whether the customer can finish their purchase. Each test type can cover a class of failure that the others don't see; concentrating everything in one type is concentrating the risk.

What we have measured on our own platform

Hábil builds a modular integration platform. In a recent measurement, across its 34 services, we counted 1,407 test classes and about 9,800 automated test methods (two different ways of counting them gave 9,771 and 9,839) [37]. With that many, one would expect to rest easy. These are lessons, now corrected, from cases in which that peace of mind was unjustified. The numbers are ours and have not been externally audited.

The payment that was made 12 times. A test of a payment flow was called "does not duplicate" and was green. It verified an internal counter of the service itself, not the call to the payment provider. When rewritten to verify the real call, it failed: 1 was expected and there were 2. With a fake server that counts how many times it is called, 12 simultaneous attempts produced 12 payments. After the fix, they produced 1. Note the difference between two kinds of double: one that answers what is expected, which hides the defect, and one that records what actually reaches it, which exposes it.

The 17 of 19 lost fields. A change in a data-extraction function passed the whole suite, which used synthetic sample documents. With real documents, the change lost 3 of 4 fields in one type of document and 17 of 19 in a foreign document. The tests were not badly written: they were fed with data that didn't resemble the real thing.

The rejection that cut off a channel. A legitimate rejection from the provider, repeated several times, tripped the circuit breaker, the mechanism that cuts off calls to a service that appears to be down. The effect was to cut off the entire channel. It was present in 7 services. The test doubles didn't repeat the rejection, so the defect had no way to show up.

Five green tests that propped up defects. We found five green tests that propped up defects, of three types: the one that blesses the error by explicit name, the one that simulates an impossible outcome and the overly permissive one. Exercising the real system ruled out 9 findings of our own and confirmed others: verifying live corrects in both directions.

Success without work. Two measurement false greens, among the most treacherous: a build that finished with the success message without having run any test, and a test run piped through a text filter that interrupted the process halfway and reported 113 and 401 tests where in fact 518 were running. The output looked like a valid count. Lesson: the number of tests is read from the report, not from a command's exit code.

The five cases share a common root: in each one, the green signal was well built and answered honestly the question it was asked. The question was the wrong one.

How do you know whether a test is any good?

There are two answers, one handcrafted and one automated.

The handcrafted one: revert the fix and see the test go red. If a test was written to protect a fix, you undo the fix and run the test: it has to fail. If it stays green, it wasn't protecting anything. For a new defect, the equivalent rule is to write the failing test first and then fix; Claude Code's official guide, for example, asks for it this way for bugs [28]. It is cheap and requires discipline. On our platform we use it as a yardstick: the 6 manual mutations we made were detected.

There is a subtler trap worth knowing about, because it doesn't appear in the literature we reviewed: a test can also pass with the defect. In one of our fixes, the criterion said "after the rejection, the balance returns to its previous value." On review we noticed that this check passed just the same with the error, because setting aside an amount and returning it leave the same number. What really distinguished them was the operations log: with the defect, two writes appeared; with the fix, none. The question worth asking of each important test is simple: would it also pass if the error were still there?

The automated one: mutation testing. Tools such as PIT (for Java) and Stryker (for JavaScript, TypeScript, C# and Scala) automatically modify the code —change a "greater than" into "greater than or equal," invert a condition, remove a call— and run the suite [18][19]. If any test fails, the change "dies"; if all pass, the change "survives" and reveals a missing test. The percentage of changes detected is called the mutation score: if the suite catches 80 of 100 sabotages, the score is 80%. PIT says it clearly: traditional coverage only measures what is executed, not whether the tests could detect a failure [18]. A study at FSE 2014 found that mutants are a valid substitute for real faults when evaluating tests [22].

Mutation has limits that a CTO should know:

  • It is slow. That is why it makes sense to apply it to the code that changed or to critical code, not to the whole repository; PIT's own documentation recommends it [18].
  • It has noise. There are "equivalent mutants": changes that don't alter behavior and that no test could detect [18]. An industrial study reported that the automatic detection of those mutants improves a lot with prior preprocessing, which indicates that the control itself has its margin of error [27].
  • It doesn't certify the business. If the test blesses a wrong rule, it can even "kill" the change that fixes the code. Research supports using mutation as a guide, not as a sole indicator: its relationship with real defects depends on the size and context of the suite [23].

That is why quality is measured as a set of signals, each with its question:

How do you know whether a test is any good?
SignalQuestion it answers
CoverageWhat code was executed?
MutationDo the checks detect plausible changes?
Verified contractsAre consumer and provider still compatible?
API and browser on critical journeysDoes the real flow work in the target environment?
Traceability between requirement and testDoes the expectation come from an approved rule, or from what the code was already doing?
Defects that reached productionDoes the evidence still predict what happens in production?

Even with real dependencies, things can still go wrong

Taking tests to real dependencies reduces one class of risk; it doesn't eliminate it. A test environment is not production: it has less data, different latency, different credentials and, often, a different version of the provider. A third party can change its behavior without notice. A load that wasn't simulated can reveal a concurrency defect. And flows that happen only once a month, such as the accounting close or the annual renewal of a policy, are rarely in the suite. This map mainly covers functional behavior and compatibility; performance, security and continuity require specific controls according to the risk.

That is why a sensible strategy has two halves. The first, more types of automated test, each aimed at a class of failure. The second, what happens after release: observing what the system does, being able to roll back and calmly reviewing what slipped through. Every defect that reaches production is information: it indicates which type of test was missing, and the right response is usually to add that test, not only to fix the code. We develop that conversation, releasing with control, in "Release without fear on your own infrastructure."

Artificial intelligence: it speeds things up, but it needs a filter

AI helps build tests: it proposes edge cases, generates skeletons, assembles API collections, explains a failure. Its output has a known profile.

What the evidence shows. An industrial study by Meta on test generation with language models measured that 75% of the cases compiled, 57% passed reliably and 73% of the recommendations were accepted by engineers [Meta study, specific sample]; all this after automatic filters that discard whatever doesn't compile, doesn't pass or adds no coverage [24]. A study on 25 JavaScript packages found a median of 70.2% statement coverage and 52.8% branch coverage with a general-purpose model [25]. These are figures from 2023 and 2024 models; today's will differ, but the direction is instructive: without filters, a significant part of what is generated is of no use.

The central risk. AI tends to copy what the code does, not what it should do. A study on 24 Java repositories concluded that tests generated with language models mostly capture current behavior, which makes it harder to detect errors [26]. It is the test that blesses the error, produced at machine speed. Other documented risks: superficial checks that inflate coverage, invented methods or rules, and sensitive data flowing out to a third party. A historical study on a code assistant found vulnerabilities in about 40% of 1,689 programs generated for specific scenarios; it is a result of its time, not a rate applicable to current models [36].

What to do about it. The vendors' guides agree on the essentials, although none brings figures of its own:

  • Give the AI a check that it can run itself; Claude Code's guide calls it the difference between a session that is watched and one that is left alone, and asks it to show the test output instead of claiming they passed [28].
  • Be specific when asking for tests: "add tests to this file" is the bad example; the good one names the edge case and what must be avoided [28].
  • Review what is generated and add the tests that are missing, as GitHub Copilot's guide indicates [30].
  • Watch out for "forced green": Anthropic's guide warns that the model may focus on making the tests pass, even with hard-coded values, and reminds us that tests are there to verify correctness, not to define the solution [29].
  • Separate whoever writes from whoever reviews. Claude Code's guide suggests different sessions for writing the tests and for writing the code that satisfies them [28].
  • Do not load production data into AI tools or test environments without classification, masking (or synthetic data) and a privacy check, and review retention and contractual terms with your provider.
  • Run what is generated through mutation: giving the AI the mutants that survived made it possible to detect up to 28% more defective pieces of human-written code than the baseline, in one study [27].

The policy that sums it all up is short: AI proposes; execution, mutation and domain review validate.

Why documenting the code improves tests

Code only says what it does. What it should do lives somewhere else, and if it isn't written down, the AI —and the new person on the team— deduces it from the implementation and copies its errors. That is where documentation comes in.

Documentation comments —Javadoc in Java, JSDoc and TSDoc in JavaScript and TypeScript, docstrings in Python— are, in practice, a contract: what a function receives, what it returns, what exceptions it can throw, what side effects it has [34]. Python's convention (PEP 257) asks for documenting behavior, arguments, return value, side effects, exceptions and usage restrictions [34]. In addition, doctest turns the examples written in the docstring into executable tests: if an example stops being true, the suite says so [34]; what the docstring explains outside those examples can indeed become outdated silently.

How much evidence is there that this improves tests? The honest answer is: it points in that direction, although we have not found a clean "same code with and without documentation" experiment. What we do have are closely related studies: converting natural-language documentation into verifiable conditions with a language model made it possible to catch 64 real historical errors [33]; extracting rules from deep learning API documentation found 94 errors versus 59 for the baseline without those constraints [33]; and a study on generating checks found improvements of 10 to 20% when incorporating Javadoc [31]. The warning goes the other way: false or outdated documentation can seriously harm a model's understanding of the code [32]. Documenting is not enough; what is documented has to be kept correct.

The third element is repository context files, such as AGENTS.md (an open format read by several coding assistants, including Codex and Copilot) or CLAUDE.md in the case of Claude Code [35]. That is where you write the commands to run the tests, the conventions and the traps that can't be deduced by reading the code. Two official recommendations: keep them short, because if they are long the assistant ignores rules, and don't let them replace approved requirements or contain secrets [28][35].

On our platform we document the code so that a new person can find their way without help,. We have not measured whether that improved the tests that assistants write, and we don't claim it. What we did measure is the cost of not writing it down: the documentation already said that a certain configuration option was not exercised outside its default value, and the gap showed up anyway. A practical rule came out of it: a configuration test must demonstrate both directions, not just the usual one.

Simple examples by industry

All are illustrative scenarios, with no customers or figures.

Simple examples by industry
IndustryScenarioWhich combination of tests catches it
BankingA transfer is duplicated when the app retries after a dropped connectionUnit test of idempotency (that repeating the same operation doesn't apply it twice: retrying doesn't repeat); integration with the real database to verify that the uniqueness constraint works; contract with the payment service; browser journey from transfer to receipt
FintechA deposit is credited twice when the provider repeats the notificationUnit test of the idempotency key and the amounts; integration with the real balance ledger; contract of the notification; API collection for the invalid signature
InsuranceThe quote applies the wrong deductible right at the age limitUnit test with the boundary ages (70 and 71, with and without a medical exam); mutation of the rating engine, because there a "greater than" swapped for "greater than or equal" changes premiums; journey from quote to issuance
RetailCheckout sells the last item to two customers at onceIntegration with the real inventory and concurrency; contract with payments; a few browser tests of the purchase journey, with the external gateway simulated
Regulated companyA regulatory report shows a wrongly aggregated totalUnit tests of the calculation rules; integration against the real schema; mutation on the aggregation logic; a requirement written in the repository context that demands a test for masking personal data in logs

Take the banking one. The unit test says the logic is correct; the integration test demonstrates that the database really rejects the duplicate; the contract test says that the payment service understands the same identifier; the browser test, that the person sees a receipt. None of the four, on its own, provides the evidence of the other three.

Where to start

  1. Measure what you already have, in two ways. Count your tests from the run report, not from the exit code, and compare it with another method. Check which active conditions your quality gate has in each project and whether they are the same.
  2. Choose your five flows with the most money or regulation at stake —a payment, an issuance, a reconciliation, a report— and ask about each one: what type of test would catch it if it breaks, and does it exist today?
  3. Review your doubles. Where the dependency decides (database, bus, provider), move from a double that answers what is expected to a real dependency or to a double that records what actually arrives.
  4. Test your tests. Choose the ten most important: revert the fix and check that they go red. In critical code, evaluate a mutation tool limited to what changes.
  5. Put AI inside the same control. Have it show the test output, not close a task with empty checks, and have each generated test backed by a written requirement.
  6. Write down what the system must do where assistants will read it: documentation of the critical rules and a short context file with the commands and the traps.

In an assessment, with an agreed scope, we can build a map of your critical flows, the existing evidence and the prioritized gaps. It does not replace a certification of the absence of defects or of compliance. How many of the tests that prop up your green today go red if the fix is reverted? That is the natural next step: the assessment can start by answering it with your own flows.

References

  1. M. Fowler, Test Pyramid (2012). https://martinfowler.com/bliki/TestPyramid.html
  2. H. Vocke, The Practical Test Pyramid (2018). https://martinfowler.com/articles/practical-test-pyramid.html
  3. K. C. Dodds, The Testing Trophy and Testing Classifications (2018). https://kentcdodds.com/blog/the-testing-trophy-and-testing-classifications
  4. Google Testing Blog, Just Say No to More End-to-End Tests (2015). One company's guidance, approximate figures. https://testing.googleblog.com/2015/04/just-say-no-to-more-end-to-end-tests.html
  5. Google Testing Blog, Where do our flaky tests come from? (2017). Google's data, not universal. https://testing.googleblog.com/2017/04/where-do-our-flaky-tests-come-from.html
  6. Google Testing Blog, Test Sizes (2010). https://testing.googleblog.com/2010/12/test-sizes.html
  7. M. Fowler, Mocks Aren't Stubs. https://martinfowler.com/articles/mocksArentStubs.html
  8. M. Fowler, Test Double. https://martinfowler.com/bliki/TestDouble.html
  9. M. Fowler, Test Coverage. https://martinfowler.com/bliki/TestCoverage.html
  10. SonarQube, Introduction to quality gates (vendor default values, configurable), consulted on Oct 6, 2026. https://docs.sonarsource.com/sonarqube-server/quality-standards-administration/managing-quality-gates/introduction-to-quality-gates
  11. SonarQube, Metrics definition (line and condition coverage), consulted on Oct 6, 2026. https://docs.sonarsource.com/sonarqube-server/user-guide/code-metrics/metrics-definition
  12. SonarQube, Test coverage overview, consulted on Oct 6, 2026. https://docs.sonarsource.com/sonarqube-server/analyzing-source-code/test-coverage/overview
  13. Testcontainers, documentation and guides, consulted on Oct 6, 2026. https://java.testcontainers.org/ · https://testcontainers.com/guides/
  14. Pact, documentation (How Pact works, What is Pact good for), consulted on Oct 6, 2026. https://docs.pact.io/ · https://docs.pact.io/getting_started/what_is_pact_good_for
  15. Postman, Test scripts and Postman CLI overview, consulted on Oct 6, 2026. https://learning.postman.com/docs/tests-and-scripts/write-scripts/test-scripts/ · https://learning.postman.com/docs/postman-cli/postman-cli-overview/
  16. Newman, official repository (README: limited maintenance; Postman CLI for new flows), consulted on Oct 6, 2026. https://github.com/postmanlabs/newman
  17. Playwright, Introduction, Best practices and Mock APIs, consulted on Oct 6, 2026. https://playwright.dev/docs/intro · https://playwright.dev/docs/best-practices · https://playwright.dev/docs/mock
  18. PIT Mutation Testing, site, mutators and FAQ, consulted on Oct 6, 2026. https://pitest.org/ · https://pitest.org/faq/
  19. Stryker Mutator, documentation, consulted on Oct 6, 2026. https://stryker-mutator.io/docs/
  20. Official documentation of JUnit, pytest, Go testing, Jest, Vitest and Testing Library, consulted on Oct 6, 2026. https://docs.junit.org/current/user-guide/ · https://docs.pytest.org/en/stable/ · https://pkg.go.dev/testing · https://jestjs.io/docs/getting-started · https://vitest.dev/guide/ · https://testing-library.com/docs/
  21. L. Inozemtseva and R. Holmes, Coverage Is Not Strongly Correlated with Test Suite Effectiveness, ICSE 2014. https://dl.acm.org/doi/10.1145/2568225.2568271
  22. R. Just et al., Are Mutants a Valid Substitute for Real Faults in Software Testing?, FSE 2014. https://dl.acm.org/doi/10.1145/2635868.2635929
  23. M. Papadakis et al., mutation and real faults, ICSE 2018 (https://doi.org/10.1145/3180155.3180183); Shin, Papadakis and Kintis, test diversity and mutation, IEEE TSE (https://doi.org/10.1109/TSE.2017.2732347).
  24. N. Alshahwan et al., Automated Unit Test Improvement using Large Language Models at Meta, FSE 2024, arXiv 2402.09171. https://arxiv.org/abs/2402.09171
  25. Schäfer et al., An Empirical Evaluation of Using Large Language Models for Automated Unit Test Generation, arXiv 2302.06527. https://arxiv.org/abs/2302.06527
  26. Study of test oracles generated with language models (24 Java repositories), arXiv 2410.21136. https://arxiv.org/abs/2410.21136
  27. Testing with language models and mutation: arXiv 2308.16557 (surviving mutants in the prompt, up to 28% more defective pieces detected) and arXiv 2501.12862 (targeted mutation in industry). https://arxiv.org/abs/2308.16557 · https://arxiv.org/abs/2501.12862
  28. Anthropic, Claude Code, Best practices and Memory, consulted on Oct 6, 2026. https://code.claude.com/docs/en/best-practices · https://code.claude.com/docs/en/memory
  29. Anthropic, Claude prompting best practices, consulted on Oct 6, 2026. https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/claude-prompting-best-practices
  30. GitHub, Copilot: write tests and Best practices, consulted on Oct 6, 2026. https://docs.github.com/en/copilot/tutorials/write-tests · https://docs.github.com/en/copilot/get-started/best-practices
  31. Liu et al., Doc2OracLL, 2025. https://doi.org/10.1145/3729354
  32. Macke and Doyle, documentation and code comprehension by language models, NAACL Findings 2024. https://aclanthology.org/2024.findings-naacl.66/
  33. Documentation as an input to tests: arXiv 2310.01831 (nl2postcond: natural language to verifiable conditions, 64 real historical errors from Defects4J) and arXiv 2109.01002 (DocTer: rules extracted from deep learning API documentation, 94 errors versus 59 for the baseline). https://arxiv.org/abs/2310.01831 · https://arxiv.org/abs/2109.01002
  34. Oracle, Javadoc doc comment specification; JSDoc; TSDoc; PEP 257; doctest documentation, consulted on Oct 6, 2026. https://docs.oracle.com/en/java/javase/21/docs/specs/javadoc/doc-comment-spec.html · https://jsdoc.app/about-getting-started · https://tsdoc.org/ · https://peps.python.org/pep-0257/ · https://docs.python.org/3/library/doctest.html
  35. AGENTS.md (open format) and Codex and GitHub Copilot documentation on repository instructions, consulted on Oct 6, 2026. https://agents.md/ · https://learn.chatgpt.com/docs/agent-configuration/agents-md · https://docs.github.com/en/copilot/reference/custom-instructions-support
  36. Pearce et al., Asleep at the Keyboard? Assessing the Security of GitHub Copilot's Code Contributions, IEEE Symposium on Security and Privacy, 2022 (https://doi.org/10.1109/SP46214.2022.9833571); version in Communications of the ACM, 2025 (https://doi.org/10.1145/3610721).
  37. Hábil, own measurements on its modular platform (34 services), 2026. In-house experience, no external audit.