Category: Artificial Intelligence

  • Progressive Delivery for AI-Generated Code: How Feature Flags Make Agentic Development Safer

    Progressive Delivery for AI-Generated Code: How Feature Flags Make Agentic Development Safer

    AI Can Write Faster Than Teams Can Safely Release

    AI coding agents are changing the tempo of software development. A team that once reviewed a few pull requests a day may now receive agent-assisted refactors, test suites, UI changes, configuration updates, and integration code in rapid succession. That speed is valuable, but it creates a new bottleneck: release confidence.

    The challenge is not only whether AI-generated code compiles, passes tests, or looks reasonable in review. The harder question is whether the change behaves safely in production, with real users, real data, real edge cases, and real business consequences.

    Progressive delivery is the release-layer discipline that helps teams answer that question gradually instead of all at once. It gives human operators a way to control who experiences a change, when they experience it, and how quickly the team can respond if something goes wrong.

    Progressive Delivery, in Plain Language

    Progressive delivery means releasing software changes in controlled steps instead of exposing every user to a new change at the same time. Teams use feature flags, staged rollouts, telemetry, and rollback plans to reduce the blast radius of mistakes.

    A feature flag is a switch in the application that lets a team turn a behavior on or off without redeploying the entire system. A staged rollout exposes a change to a small group first, then expands access if metrics and feedback look healthy. A kill switch is a preplanned way to quickly disable a risky capability when something goes wrong.

    • Traditional release: merge code, deploy it, and every user gets the change at once.
    • Progressive release: merge code, deploy it safely, keep it hidden or limited, then expand exposure based on evidence.
    • AI-first release: treat code, prompts, model choices, configuration, and agent behaviors as releasable artifacts that need controls.

    Deployment and Release Should Not Be the Same Event

    For AI-assisted teams, one of the most important mental shifts is separating deployment from release. Deployment means the code or configuration is available in an environment. Release means users can actually experience the change.

    When deployment and release are tied together, every deployment becomes a high-stakes event. When they are decoupled, teams can deploy more often while releasing more carefully. A risky feature can be deployed "dark," meaning it exists in production but is not visible to most users. Engineers can test it internally, enable it for a narrow cohort, watch telemetry, and expand only when the evidence supports it.

    This matters even more when code is generated or heavily modified by AI agents. Agents can produce implementation detail quickly, but they do not automatically understand every product constraint, customer expectation, compliance requirement, or operational nuance. Progressive delivery gives humans a control plane for deciding when generated work should reach users.

    A Safer Rollout Workflow for Agent-Generated Changes

    Progressive delivery works best when it is treated as a repeatable workflow, not as a last-minute safety net. A practical AI-first release process can follow these steps:

    • Classify the risk: Label the change as low, medium, or high risk based on user impact, data sensitivity, reversibility, and operational complexity.
    • Assign an owner: Make one person or team accountable for the rollout plan, monitoring, rollback decision, and cleanup.
    • Wrap risky behavior in a flag: Put new logic, UI paths, automation, or AI behaviors behind a feature flag before deployment.
    • Deploy dark: Ship the code to production with the flag off for general users, then verify that the application remains stable.
    • Test internally: Enable the flag for developers, QA, support staff, or a small internal group before exposing it externally.
    • Roll out gradually: Expand by cohort, account type, geography, percentage, allowlist, or other meaningful segment instead of turning the feature on for everyone.
    • Monitor release telemetry: Watch error rates, latency, conversion, task completion, support tickets, model cost, token usage, and user feedback.
    • Pause, expand, or roll back: Make release decisions based on agreed thresholds, not optimism or pressure to ship.
    • Remove the flag when done: Once a change is fully released and stable, schedule cleanup so temporary release controls do not become permanent clutter.

    What Belongs Behind a Flag?

    Teams often associate feature flags with visible UI changes, but AI-first software broadens the list. If a change can affect user experience, cost, trust, safety, or data behavior, it may deserve a controlled release path.

    • User interface changes: New layouts, navigation updates, onboarding flows, dashboards, or editor experiences.
    • Pricing and workflow logic: Plan limits, checkout behavior, usage caps, upgrade prompts, entitlement checks, or approval flows.
    • Database and migration behavior: New write paths, backfills, schema-dependent logic, or data transformation jobs that can be enabled gradually.
    • AI prompt changes: Updated system prompts, retrieval instructions, tone rules, summarization formats, or escalation criteria.
    • Model switches: Moving from one model to another, changing model parameters, or routing different cohorts to different inference providers.
    • Autonomous-agent behavior: New tool permissions, background tasks, lead research flows, content generation steps, or automated remediation actions.
    • WordPress plugin behavior: For an AI content pipeline, chat assistant, or CRM-style lead research plugin, staged rollout thinking could apply to generated post workflows, chat escalation logic, or automated prospecting steps.

    The key question is simple: if this change behaves badly, how quickly can we limit harm? If the answer is "not quickly," the change probably needs a flag, a rollout plan, a kill switch, or all three.

    Telemetry Turns Rollouts Into Decisions

    Feature flags are most powerful when they are connected to telemetry. Without measurement, a staged rollout can become a slower version of guessing. With measurement, teams can define what healthy release behavior looks like before they expand exposure.

    Useful rollout metrics depend on the change, but common examples include application error rate, API latency, failed jobs, conversion rate, user task completion, support contacts, cancellation signals, token spend, hallucination reports, moderation events, and manual override frequency. For agentic features, teams should also monitor tool-call failures, escalation rates, retry loops, and unexpected output patterns.

    The release owner should know which metric would cause an immediate pause, which metric would trigger rollback, and which metric would justify expanding from 5 percent to 25 percent to 100 percent. This keeps rollout decisions grounded in evidence rather than enthusiasm.

    Kill Switches Are Not a Sign of Failure

    A kill switch is a sign that the team planned responsibly. It gives operators a fast, low-drama way to disable a risky capability without waiting for a new build, emergency deploy, or late-night debugging session.

    Good kill switches are specific enough to avoid unnecessary disruption. Instead of shutting down an entire product, a team may disable only a new recommendation model, a background agent task, a migration worker, a chat escalation path, or an experimental checkout rule. The goal is to contain the blast radius while keeping the rest of the system useful.

    The Tradeoffs: Progressive Delivery Is Not Free

    Progressive delivery adds safety, but it also adds operational complexity. Teams should adopt it deliberately and manage the costs rather than assuming flags automatically make every release safe.

    • Flag debt: Old flags accumulate and make code harder to understand unless ownership and cleanup dates are assigned.
    • Testing complexity: Every flag can create multiple application states, so teams need a sensible strategy for testing important combinations.
    • Inconsistent user experiences: Staged rollouts can mean different users see different behavior, which may complicate support, documentation, and sales conversations.
    • False confidence: Automation can detect many failures, but it cannot replace product judgment, customer empathy, security review, or incident preparedness.
    • Governance overhead: High-risk flags need clear approval, auditability, and access control so release switches do not become informal production backdoors.

    A Practical Checklist for AI-First Teams

    Progressive delivery does not need to begin as a large platform initiative. A small team can start with a few habits that make releases safer immediately.

    • For founders: Identify product behaviors that could harm trust, revenue, or customer operations if they changed unexpectedly. Require flags or kill switches for those areas.
    • For engineering managers: Define release ownership. Every flagged change should have an owner, rollout plan, rollback plan, success metrics, and cleanup date.
    • For developers: Add flags before merging risky work, not after a scare. Keep flag names clear, document intended removal, and test both enabled and disabled states.
    • For AI workflow owners: Treat prompts, model selections, agent permissions, and configuration changes as release artifacts. Review and roll them out with the same care as code.
    • For support and operations: Know which flags affect user-facing behavior and where to report unusual patterns during staged releases.
    • For everyone: Decide in advance what "stop," "pause," and "expand" mean for each rollout. The middle of an incident is the worst time to invent the rules.

    The Industry Is Moving Toward Governed, Observable Releases

    The broader software tooling market is moving in the same direction: more AI assistance, more automation, and more need for release control. Cloudflare introduced Flagship as a feature flag platform built for AI-era development. Datadog launched Feature Flags to connect rollout decisions with observability. AWS published guidance on feature flag orchestration with AWS DevOps Agent and LaunchDarkly. Atlassian has described an AI-enabled software development lifecycle, while GitLab has announced capabilities focused on giving enterprises speed and control at scale.

    Sources and Fact Check References

    • Cloudflare Blog – Cloudflare introduced Flagship as a feature flag platform built for AI-era development.
    • Datadog – Datadog launched Feature Flags to connect rollout decisions with observability.
    • AWS DevOps Blog – AWS published guidance on feature flag orchestration with AWS DevOps Agent and LaunchDarkly.
    • Atlassian – Atlassian has described an AI-enabled software development lifecycle.
    • GitLab Blog – GitLab has announced capabilities focused on giving enterprises speed and control at scale.
  • Why MCP Matters: Giving AI Coding Agents Safe Access to Your Tools and Data

    Why MCP Matters: Giving AI Coding Agents Safe Access to Your Tools and Data

    The Problem: Smart Assistants, Disconnected Workflows

    AI coding assistants are now useful for explaining code, drafting functions, generating tests, and suggesting fixes. But many still work from a narrow view of the project: the prompt you typed, the files you opened, and perhaps a recent repository snapshot.

    Real software development is broader than that. A useful agent may need to inspect a GitHub issue, read internal documentation, check CI status, review a feature flag, consult product requirements, or compare behavior against a database record. Without those connections, the assistant can sound confident while missing the context that actually determines the right answer.

    That is why Model Context Protocol, usually shortened to MCP, matters. MCP is not just another AI trend label. It is a concrete integration pattern for connecting large language model applications and agents to the tools and data sources teams already use. In AI-first development, that integration layer may become as important as the editor, the issue tracker, or the CI pipeline.

    MCP in Plain Language

    MCP is an open protocol that lets AI applications connect to external tools, data sources, and reusable context through a common interface. Instead of every coding assistant needing a custom integration for every database, documentation system, ticket tracker, or internal API, MCP defines a shared way for those systems to expose capabilities to an AI host.

    A common analogy is USB-C for AI context. The point is not that every connected system is identical. The point is that there is a standard way to connect, discover what is available, request an action, and return results. For software teams, that can reduce one-off glue code and make integrations easier to reuse, review, and govern.

    The Basic MCP Mental Model

    An MCP setup usually includes a host application, an MCP client, and one or more MCP servers. The host is the AI application the user interacts with, such as a coding environment or AI desktop assistant. The client manages the connection between that host and a server. The MCP server exposes specific capabilities from an external system, such as a repository, documentation index, database, project tracker, browser automation layer, or internal service.

    • Tools are callable actions, such as searching issues, checking build status, creating a draft pull request, or querying a read-only database view.
    • Resources are structured pieces of context the agent can read, such as files, documentation pages, logs, design notes, or product requirements.
    • Prompts are reusable interaction templates that can guide a model through a known workflow, such as triaging a bug report or summarizing a release plan.
    • Permissions define what the host and user allow the agent to access or do. Good MCP usage should make capabilities explicit rather than hiding them inside vague automation.
    • Auditability means tool calls, inputs, outputs, and approvals should be visible enough for humans to understand what happened and why.

    That last point is essential. MCP makes agents more capable, but capability is not the same as safety. A coding agent that can read a README is low risk. A coding agent that can modify production data, rotate secrets, merge pull requests, or email customers is a very different kind of system.

    Practical Examples in Software Development

    The practical value of MCP appears when a coding agent can combine code context with workflow context. Imagine asking an agent, "Why is this checkout test failing?" Without tool access, it may only inspect the test and make an educated guess. With carefully scoped MCP servers, it could review the related issue, inspect recent pull requests, check the CI failure log, search internal docs for payment provider behavior, and propose a targeted fix.

    • Issue triage: The agent reads a GitHub issue, identifies the affected package, checks linked discussions, and proposes reproduction steps.
    • Documentation lookup: The agent searches team docs or API references before changing code, reducing guesswork and hallucinated interfaces.
    • CI awareness: The agent checks failing jobs, summarizes the first meaningful error, and suggests whether the issue is test flakiness, configuration drift, or a real regression.
    • Pull request drafting: The agent prepares a draft PR description, links relevant issues, lists risk areas, and flags tests that should be reviewed by a human.
    • Product requirement review: The agent compares a proposed implementation against a product brief or acceptance criteria before touching code.

    These examples matter because they connect the agent to the work system, not just the codebase. In many teams, the truth is distributed across tickets, docs, logs, dashboards, tests, and conversations. MCP gives AI tools a more consistent path into that distributed context.

    How MCP Differs from Plugins, Scripts, and Direct APIs

    Teams have always connected tools with scripts and APIs. A developer can write a bot that calls GitHub, reads a database, posts to Slack, and updates a ticket. That can work well for a narrow workflow. The weakness is that each integration often invents its own conventions for authentication, schema design, error handling, prompts, and permissions.

    One-off plugins have a similar limitation. They may be convenient, but they are often tied to one vendor, one host application, or one workflow. MCP's promise is a more portable integration model: build or approve an MCP server once, then connect it to compatible AI hosts under explicit controls. That does not eliminate engineering work, but it can reduce duplication and make governance easier.

    Direct API integrations still matter, especially for production-grade systems with strict performance, compliance, or reliability requirements. MCP is better understood as an agent-facing integration layer. It helps AI tools discover and use capabilities in a structured way. It does not replace thoughtful API design, secure infrastructure, or application-level authorization.

    The Tradeoff: More Context, More Risk

    Disconnected assistants are limited. Connected agents are powerful. That power creates a larger risk surface. The central operational question is not "Can we connect this tool?" but "What should the agent be allowed to see or do, under which conditions, and with what human oversight?"

    • Security exposure: Every server, token, and connected system can become a path to sensitive data or unsafe actions.
    • Permission sprawl: Teams may start with a few safe read-only tools and slowly accumulate broad access that no one actively reviews.
    • Prompt-injection risk: If an agent reads untrusted content from issues, web pages, documents, or customer messages, that content may try to manipulate the agent's behavior.
    • Brittle tool schemas: Poorly described tools can cause agents to call the wrong action, misunderstand parameters, or treat partial results as complete truth.
    • Over-automation: Just because an agent can open, edit, merge, deploy, or notify does not mean it should do so without human approval.

    The healthiest teams will treat MCP servers like part of their software supply chain. Servers should be reviewed, versioned, documented, monitored, and retired when they are no longer needed. Convenience is valuable, but invisible convenience is dangerous.

    A Starter Checklist for Small Teams

    Small teams do not need an enterprise governance program to use MCP responsibly. They do need clear defaults. A practical starting point is to make the first integrations boring, read-only, and easy to observe.

    • Begin read-only. Start with documentation search, issue lookup, CI log reading, or repository inspection before enabling write actions.
    • Use least privilege. Give each MCP server only the access required for its specific job, not a broad personal token with sweeping permissions.
    • Separate dev, staging, and production. An agent that can experiment in development should not automatically have production access.
    • Log tool calls. Keep records of what the agent called, what inputs it sent, what came back, and which user approved the action.
    • Review server provenance. Know who built the MCP server, how it is maintained, what dependencies it uses, and whether it handles secrets safely.
    • Document approved servers. Maintain a simple internal list of allowed MCP servers, owners, scopes, and acceptable use cases.
    • Require human approval for destructive actions. Deleting data, merging code, changing permissions, sending external messages, or triggering deployments should remain gated.

    This checklist is intentionally conservative. The goal is not to slow teams down forever. The goal is to earn trust step by step, so automation expands only where it has proven useful and controllable.

    Why This Matters Beyond the Code Editor

    MCP is especially relevant for AI-first product workflows because useful automation rarely lives in one system. An AI-assisted publishing pipeline may need scoped access to drafts, editorial rules, schedules, and content history. A website chat assistant may need visitor context, support status, escalation rules, and knowledge base entries. A CRM lead workflow may need to consult public records, enrich a lead profile, and record why a suggestion was made.

    In WordPress and product environments, the same rule applies: the agent should get the context it needs, but not unlimited access to everything the site or business knows. A publishing assistant does not need billing permissions. A chat assistant does not need broad database write access beyond its support workflow. A lead research agent should record sources and respect limits on what it can collect or change.

    What to Watch as MCP Matures

    MCP's future will depend on more than technical elegance. Adoption will be shaped by server quality, permission design, registry trust signals, enterprise policy support, and how clearly hosts present tool activity to humans. If the experience is too permissive, teams will block it. If it is too clumsy, developers will bypass it with scripts. The winning pattern is likely to be structured, observable, and boring in the best sense of the word.

    AI coding agents are already moving toward more agentic workflows, where they can plan tasks, inspect context, run commands, and propose changes. MCP helps make those connections more explicit and reusable. For teams adopting AI-first development, the opportunity is not just faster code generation. It is better-connected workflows with clearer boundaries, stronger review habits, and safer paths from idea to implementation.

    Sources and Fact Check References

  • The New Code Review: How Humans Should Review Work From AI Coding Agents

    The New Code Review: How Humans Should Review Work From AI Coding Agents

    AI Can Write the Diff. Humans Still Own the Decision.

    AI coding agents are changing what code review is for. In a traditional review, a teammate usually explains the problem, writes the code, and opens a pull request with human intent behind every major choice. With an AI coding agent, implementation can arrive faster, broader, and sometimes more confidently than the underlying reasoning deserves.

    That does not make review less important. It makes review more judgment-heavy. The reviewer’s job is no longer just to spot syntax mistakes, suggest cleaner names, or ask for one more test. It is to decide whether the change should exist, whether it solves the right problem, whether it fits the system, and whether the team can safely maintain it later.

    Recent industry research points in the same direction: AI adoption in software work is rising, but trust, accuracy, and human verification remain central concerns. Stack Overflow’s 2025 Developer Survey found that more developers distrusted the accuracy of AI tools than trusted it, while DORA’s 2025 research reported broad workplace use of AI among technology professionals alongside ongoing questions about effective, reliable adoption. In practice, strong teams treat AI-generated code as a fast draft from a capable but non-accountable contributor. Useful? Often. Final? Not until a human has reviewed it.

    Why AI-Written Code Needs a Different Review Mindset

    AI coding agents are good at producing plausible code. That is both their strength and their risk. A human junior developer may ask clarifying questions, hesitate around unfamiliar systems, or leave obvious gaps. An AI agent may produce a complete-looking implementation even when the task is underspecified, the repository patterns are unclear, or the business rule is ambiguous.

    Reviewers should assume three things until proven otherwise: the agent may have optimized for local correctness instead of system fit, it may have filled in missing requirements without saying so, and it may have changed more than the task required. This is not a reason to reject AI assistance. It is a reason to review from the outside in.

    • Do not start by admiring the diff. Start by restating the user need or engineering goal.
    • Do not assume a passing test means the behavior is right. Ask whether the test proves the intended outcome.
    • Do not treat confident code as explained code. Require traceable reasoning for important changes.
    • Do not reward large, sweeping changes if a smaller change would have solved the problem.
    • Do not let the AI agent’s speed pressure the team into lowering review standards.

    Before Reading the Diff, Check the Assignment

    The most useful review often happens before the reviewer opens the changed files. If the task is vague, the code review will become a guessing game. For AI-generated work, reviewers should first inspect the prompt, ticket, acceptance criteria, or issue description that guided the agent.

    Ask whether the agent was given a clear target. What behavior should change? What should stay the same? Which files, APIs, roles, devices, permissions, or data boundaries matter? What constraints were stated? What constraints were assumed? If the task asks for “improve checkout validation,” the reviewer needs to know whether that means better error messages, stricter server-side rules, accessibility improvements, fraud prevention, or all of the above.

    • What exact problem is this change supposed to solve?
    • Who benefits from the change: user, admin, developer, support team, or business stakeholder?
    • What are the acceptance criteria, and are they measurable?
    • What areas of the system were intentionally out of scope?
    • Was the AI agent allowed to add dependencies, change database schemas, alter public APIs, or refactor unrelated code?
    • Is there a human-readable summary of what the agent changed and why?

    A Layered Review Workflow for AI Coding Agents

    A practical human-in-the-loop review works best in layers. Instead of reading every line from top to bottom immediately, move from purpose to risk to implementation detail. This helps reviewers avoid getting distracted by polished code that may not solve the right problem.

    1. Product Intent: Does This Solve the Right Problem?

    Start with the outcome. If the change is user-facing, verify that it matches the intended workflow, language, permission model, and failure states. If it is internal, verify that it improves the developer or operational experience without creating hidden obligations.

    AI agents can accidentally implement a nearby idea instead of the actual requirement. For example, an agent asked to “add admin filtering” might build a new search interface when the real need was a simple status dropdown on an existing table. The code may work, but the product judgment is wrong.

    • Does the change match the original request, not merely a related interpretation?
    • Are edge cases defined from the user’s point of view?
    • Could the new behavior surprise existing users?
    • Are copy, labels, errors, and empty states clear and appropriate?
    • Does the change respect role permissions and business rules?

    2. Architecture Fit: Does It Belong Here?

    Next, check whether the implementation fits the existing system. AI agents often infer patterns from nearby files, but they may miss deeper conventions: service boundaries, domain ownership, performance assumptions, release constraints, or framework-specific best practices.

    A good reviewer asks whether the change makes the codebase easier or harder to reason about six months from now. A solution that adds a new abstraction, helper, dependency, or background job should justify the extra moving parts.

    • Does the change follow existing project patterns?
    • Is the logic located in the right layer, such as UI, API, domain service, or data access?
    • Does it duplicate behavior that already exists elsewhere?
    • Does it introduce a new abstraction before the codebase needs one?
    • Would another developer know where to look when this feature breaks?

    3. Data, Security, and Privacy Risk: What Could Go Wrong?

    AI-generated code deserves careful review anywhere it touches authentication, authorization, payments, personally identifiable information, customer data, logs, file uploads, external APIs, or database writes. These are areas where a small plausible mistake can become a serious incident.

    Reviewers should pay special attention to silent trust changes. Did the code move validation from the server to the client? Did it expose extra fields in an API response? Did it log sensitive input? Did it make an admin-only operation reachable from a lower-privilege path? These problems may not stand out in a diff unless the reviewer is looking for them.

    • Are authorization checks still enforced on the server?
    • Are inputs validated and outputs encoded in the right places?
    • Does the change expose new data through responses, logs, analytics, or error messages?
    • Are secrets, tokens, and credentials handled safely?
    • Do database migrations preserve existing data and support rollback?
    • Does any new dependency increase supply-chain risk?

    4. Test Evidence: What Has Been Proven?

    For AI-generated work, reviewers should not ask only “Are there tests?” A better question is “What claim do these tests prove?” AI agents can create tests that mirror their own assumptions, assert implementation details, or cover the happy path while missing the real failure mode.

    Useful tests connect back to acceptance criteria. If the task is about permissions, tests should cover allowed and denied users. If the task is about data transformation, tests should include messy inputs. If the task is about a user interface, tests or review evidence should cover keyboard navigation, screen states, and error handling where appropriate.

    • Do the tests fail without the production change?
    • Do they cover the bug, feature, or risk described in the task?
    • Are negative cases included, not only happy paths?
    • Are edge cases represented with realistic data?
    • Is there evidence from local runs, CI, screenshots, logs, or manual verification when automated coverage is not enough?

    5. Readability and Maintainability: Can Humans Own This Code?

    AI agents can generate code that is syntactically correct but oddly shaped. The reviewer should make sure future humans can understand, debug, and extend it. Cleverness is not a virtue if it makes the team dependent on another AI pass to understand the implementation.

    Look for unnecessary generalization, inconsistent naming, overly defensive branches, and comments that describe what the code does without explaining why. Also watch for large formatting churn that hides the meaningful change.

    • Is the simplest reasonable solution used?
    • Are names consistent with the domain language of the project?
    • Can the code be understood without reading the original prompt?
    • Are comments used to explain non-obvious decisions rather than restating the code?
    • Does the diff avoid unrelated cleanup, formatting churn, and opportunistic refactors?

    6. Operational Impact: What Happens After Merge?

    Some changes are correct in isolation but risky in production. Reviewers should consider deployment, monitoring, performance, support, and rollback. AI agents may not know which parts of the system are fragile, expensive, rate-limited, or heavily used unless the prompt and repository context made that clear.

    • Could this increase latency, memory use, API calls, database load, or background job volume?
    • Does the change need feature flags, staged rollout, or migration sequencing?
    • Are errors observable through logs, metrics, or alerts?
    • Can the change be rolled back safely?
    • Will support, documentation, or customer-facing guidance need updates?

    When to Ask the AI Agent for a Self-Review

    A useful habit is to ask the AI coding agent to review its own work before the human review begins. This is not a substitute for human judgment. It is a way to surface assumptions, summarize changes, and generate a checklist of likely risk areas.

    Sources and Fact Check References

    • Stack Overflow Developer Survey 2025 – Stack Overflow’s 2025 Developer Survey found that more developers distrusted the accuracy of AI tools than trusted it.
    • DORA 2025 Research – DORA’s 2025 research reported broad workplace use of AI among technology professionals and examined reliable adoption of AI in software delivery.
  • When AI Writes the Tests: How to Keep Agentic Development Fast Without Letting Flaky Checks Slow You Down

    When AI Writes the Tests: How to Keep Agentic Development Fast Without Letting Flaky Checks Slow You Down

    The New Bottleneck: Trusting the Tests

    Picture a small product team preparing a release. An AI coding agent has implemented a feature, updated a few files, and helpfully generated new tests. The pull request looks impressive: coverage is higher, the suite passes, and the change appears ready before lunch. Then the same tests fail on the next run with no code changes. Or worse, they keep passing while a real bug slips into production.

    That is the tension of agentic development. AI tools can speed up more than application code. They now write tests, update mocks, propose CI configuration, add fixtures, revise release scripts, and summarize changes for reviewers. The question is no longer whether AI can help create tests. It is whether those tests are trustworthy enough to protect the product.

    Flaky Tests and Quality Gates, in Plain Language

    A flaky test is a test that sometimes passes and sometimes fails without a meaningful change in the software being tested. Flakiness can come from timing assumptions, random data, shared state, external network calls, file-system differences, time zones, race conditions, or tests that depend on being run in a particular order.

    A quality gate is a rule in the delivery pipeline that decides whether a change is allowed to move forward. Common quality gates include passing unit tests, minimum coverage thresholds, static analysis checks, security scans, required code review, and deployment approvals. In healthy CI/CD, quality gates are not bureaucracy. They are the automated and human checkpoints that help fast teams avoid preventable problems.

    When AI agents generate tests, the quality gate itself needs scrutiny. A test suite that passes is useful only if it checks the right behavior in a repeatable way.

    Why AI Agents Are Touching More Than Application Code

    Modern coding agents are increasingly used as end-to-end development assistants. A developer may ask an agent to fix a bug, and the agent may respond by editing source code, adding a regression test, updating snapshots, modifying CI commands, and summarizing the change. That is useful because real software work is rarely limited to one file.

    Industry research points toward broader adoption of coding agents across development workflows. Anthropic’s 2026 Agentic Coding Trends Report describes software development as shifting toward orchestrating agents that write code and highlights ongoing tradeoffs around productivity, oversight, quality, and security. Gartner also reported in May 2026 that the enterprise AI coding agent market was entering a phase of expansion and competitive realignment.

    The benefit is obvious: AI can produce first-draft tests faster than most teams can write them by hand. The risk is quieter: agents are often optimized to satisfy the visible request. If the prompt says, “add tests and make CI pass,” an agent may write tests that are technically valid but weak, over-mocked, too tightly coupled to implementation details, or blind to the behavior users actually depend on.

    A Good Test Is More Than a Test That Exists

    A good test protects an important behavior. It should fail when that behavior breaks and pass when the behavior works. That sounds simple, but it is exactly where many AI-generated tests need human review.

    A weak test might verify that a function was called instead of verifying the result users care about. A brittle test might assert the exact wording of an internal error message that was never part of the product contract. An over-mocked test might replace every dependency with fake objects, proving only that the mocks behave as expected. A snapshot test might lock in a large block of output without making clear which part matters.

    Good tests tend to be specific, deterministic, readable, and connected to real risk. They explain the system’s expected behavior in a way another developer can understand six months later. AI can help draft them, but engineering judgment decides whether they are meaningful.

    Common Failure Modes in Agent-Generated Tests

    • Brittle assertions: The test checks incidental details, such as private method calls, object ordering that is not guaranteed, or exact formatting that users never see.
    • Excessive mocking: The test replaces so much of the system that the meaningful integration path is never exercised.
    • False confidence: Coverage increases, but the new tests do not check edge cases, failure handling, permissions, data integrity, or user-visible outcomes.
    • Nondeterministic behavior: The test depends on real time, random values, network availability, file-system state, local configuration, or test execution order.
    • Fixture sprawl: The agent creates large test fixtures that are hard to understand, hard to maintain, and easy to accidentally misuse.
    • Snapshot overload: The test approves a large generated output without explaining which fields are important and which are incidental.
    • Happy-path bias: The test confirms the ideal case but ignores invalid input, empty states, rate limits, authentication boundaries, and recovery from failed dependencies.
    • CI mismatch: The test passes locally but fails in CI because the agent assumed a different runtime, database state, environment variable, locale, or dependency version.

    Recent research into agent-generated tests reinforces the point. A July 2026 arXiv paper analyzing 204,673 test artifacts from the AIDev dataset reported that agent-generated tests showed stronger edge-case variety than human-authored tests in the studied sample, but also a higher candidate rate for flakiness, largely tied to file I/O and nondeterministic logic. In other words, AI-written tests can be useful and still require review for robustness.

    A Lightweight Review Checklist Before Merging

    Teams do not need a heavyweight process for every AI-generated test. They do need a consistent review habit. Before merging a pull request that contains agent-written or agent-modified tests, ask these questions:

    • What behavior is this test protecting? If the answer is not clear, rename or rewrite the test.
    • Would this test fail if the real bug came back? Regression tests should prove the fix, not just execute nearby code.
    • Is the test deterministic? Remove dependence on real time, random data, network calls, shared files, or execution order unless those are deliberately controlled.
    • Are the mocks hiding the risk? Mock external systems where necessary, but keep enough real behavior to validate the integration that matters.
    • Is the assertion about an outcome or an implementation detail? Prefer user-visible results, persisted state, emitted events, API responses, or documented contracts.
    • Is the fixture small and intentional? Test data should be readable and relevant, not a large blob created just to satisfy setup requirements.
    • Does the test cover failure paths? AI often writes happy-path tests first; reviewers should look for permissions, invalid input, empty data, retries, and error handling.
    • Will this test be understandable later? If a future maintainer cannot tell why it exists, it is not finished.

    CI/CD Guardrails That Keep Speed From Becoming Chaos

    Quality gates work best when they make the desired behavior easy and risky behavior visible. For AI-generated tests, the goal is not to slow teams down. The goal is to prevent a fast feedback loop from becoming a noisy feedback loop.

    • Use deterministic fixtures: Keep test data stable, minimal, and isolated. Seed databases predictably and avoid depending on production-like randomness.
    • Isolate test data: Each test should create and clean up its own data or run inside a disposable environment. Shared state is a common source of flakiness.
    • Block real network calls by default: Unit tests and most integration tests should not depend on live third-party services. Use recorded responses, contract tests, or controlled test doubles.
    • Control time and randomness: Freeze clocks, seed random generators, and avoid tests that change behavior based on the current date or local time zone.
    • Set coverage thresholds carefully: Coverage can prevent backsliding, but it should not reward meaningless tests. Use it as one signal, not the only signal.
    • Consider mutation testing where appropriate: Mutation testing can reveal whether tests actually detect changed behavior, though it may be too slow or costly for every pipeline.
    • Require human review for high-risk paths: Authentication, payments, data deletion, privacy-sensitive workflows, migrations, and permission logic deserve explicit human approval.
    • Add flaky-test quarantine policies: If a test is flaky, track it, quarantine it temporarily if needed, assign ownership, and fix or delete it. Do not let random failures become normal.
    • Measure test health over time: Track retry rates, duration changes, failure frequency, and which tests are most often quarantined. Observability applies to the test suite too.

    A practical pipeline might run fast deterministic tests on every pull request, deeper integration tests before merge, and slower end-to-end or mutation checks on a schedule. Not every repository needs the same gates. A small plugin team and a large enterprise platform will make different tradeoffs, but both need confidence that passing CI means something.

    CI/CD itself is also becoming part of the agentic surface area. A 2026 arXiv study of 8,031 agentic pull requests across 1,605 GitHub repositories found that CI/CD configuration files accounted for 3.25% of agent changes, with most of those changes targeting GitHub Actions. That makes pipeline review part of the same quality conversation as test review.

    What This Means for WordPress and AI Plugin Teams

    WordPress teams building AI-enabled products face a particularly interesting version of this problem. Plugins often interact with databases, scheduled jobs, user roles, REST APIs, admin screens, external AI services, and third-party themes or plugins. That creates many places where an AI-generated test can look convincing while missing the real integration risk.

    For example, a team building an AI pipeline plugin for scheduled publishing, a chat assistant that escalates to a human, or a CRM plugin that enriches lead records should care about regression checks around permissions, rate limits, data persistence, cron behavior, and failure recovery. In a context like CoatiPress, reliable tests would not just confirm that an AI call was mocked successfully.

    Sources and Fact Check References

    • Anthropic – Anthropic’s 2026 Agentic Coding Trends Report describes software development as shifting toward orchestrating agents and discusses productivity, oversight, quality, and security tradeoffs.
    • Gartner – Gartner reported in May 2026 that the enterprise AI coding agent market was entering a phase of expansion and competitive realignment.
    • arXiv – A July 2026 arXiv paper analyzed 204,673 test artifacts from the AIDev dataset and reported higher candidate flakiness in agent-generated tests, largely tied to file I/O and nondeterministic logic.
    • arXiv – A 2026 arXiv study of 8,031 agentic pull requests across 1,605 GitHub repositories found that CI/CD configuration files accounted for 3.25% of agent changes, with most targeting GitHub Actions.
  • From Copilot to Coding Agents: How AI-First Development Is Changing the Pull Request

    From Copilot to Coding Agents: How AI-First Development Is Changing the Pull Request

    Why coding agents suddenly feel more real

    For years, AI in software development mostly meant autocomplete: a helpful suggestion inside the editor, a generated function, or a chat answer explaining an error message. That kind of assistance is still useful, but the bigger shift is toward agentic development workflows. These tools can read more of a repository, form a plan, edit multiple files, run tests when permitted, respond to failures, and prepare changes for a human to review.

    That does not mean teams should hand production systems to an AI and hope for the best. It means the unit of work is changing. Instead of asking, “Can AI write this line?” teams are asking, “Can AI take this scoped issue, work in a branch, follow our project rules, pass checks, and produce something reviewable?” That is the heart of AI-first development: turning intent, context, and verification into a repeatable workflow.

    Completion, chat, and agents are not the same thing

    The phrase “AI coding tool” now covers several different workflows. Separating them helps teams set realistic expectations and choose the right level of autonomy for each task.

    • Code completion suggests snippets as a developer types. It is fast, local to the current file, and best for boilerplate, common patterns, and small transformations.
    • Chat-assisted coding lets a developer ask questions, paste errors, request explanations, or generate code through back-and-forth guidance. It is useful for learning, debugging, and exploring options, but the human usually drives each step.
    • Agentic coding workflows assign a bounded task to an AI system that can inspect broader project context, make changes across files, run approved commands or tests, and return a proposed diff or pull request. The human shifts from typing every edit to specifying intent, reviewing results, and enforcing quality.

    The difference is more than interface design. A completion tool lives in the moment of writing. A coding agent can operate around an issue, branch, test run, or pull request. That makes it powerful, but it also makes guardrails more important.

    How today’s coding-agent workflows compare

    The leading tools are converging on a similar idea: give the model enough repository context and a bounded task, then let it produce reviewable work. They differ in where they live, how asynchronous they are, and how much control they give teams over environment, permissions, and review.

    • GitHub Copilot agent mode is designed around GitHub and editor-based workflows. GitHub describes agent mode as enabling Copilot to iterate on its own output, fix errors, suggest terminal commands, and analyze run-time errors in pursuit of a user’s request.
    • OpenAI Codex is positioned as a software engineering coding agent for real engineering work, including routine pull requests, features, refactors, migrations, testing, code review, and background tasks.
    • Google Jules emphasizes asynchronous agent work: developers can connect a repository, choose a branch, submit a task, review a generated plan, and come back when the work completes or needs input.
    • Claude Code focuses on terminal and repository workflows, with best practices around giving the agent clear context, asking it to plan, iterating through tests, and applying project-specific instructions.

    There is no universal winner for every team. A startup building quickly, an enterprise with strict compliance needs, a WordPress plugin shop, and an open-source maintainer may all value different capabilities. The practical question is not “Which agent replaces developers?” It is “Which workflow fits our repo structure, testing culture, review process, and risk tolerance?”

    What coding agents are good at today

    Coding agents are most useful when the task is concrete, the expected outcome is easy to verify, and the repository contains enough patterns for the agent to follow. They are less reliable when requirements are vague, domain context is missing, or success depends on product judgment rather than technical execution.

    • Drafting or updating documentation based on existing code and configuration.
    • Writing first-pass unit tests for functions, classes, API endpoints, and known edge cases.
    • Fixing small bugs with clear reproduction steps and failing tests.
    • Applying dependency updates, lint fixes, formatting changes, and repetitive migrations.
    • Refactoring narrow areas of code while preserving existing behavior.
    • Explaining unfamiliar modules to new team members or technical leaders.
    • Preparing pull request summaries that describe changed files, risks, and test coverage.

    These strengths map well to work that many teams postpone because it is necessary but time-consuming. A coding agent that drafts tests, updates docs, or handles a small bug can create leverage without asking the organization to trust it with major architectural decisions.

    What still needs human review

    Human judgment remains central. AI can produce code that looks plausible while missing an edge case, misunderstanding a requirement, or introducing a security issue. Review is not a formality; it is where engineering responsibility stays with the team.

    • Product intent: Does the change solve the right problem for real users?
    • Architecture: Does it fit the system’s long-term design, or does it add hidden complexity?
    • Security: Does it validate input, escape output, protect secrets, and respect permission boundaries?
    • Performance: Does it introduce slow queries, unnecessary network calls, or expensive loops?
    • Maintainability: Will the next developer understand the change six months from now?
    • Release risk: Can the team roll back safely if the change behaves unexpectedly?

    A useful mental model is to treat an AI agent like a very fast junior contributor with unusual memory and no lived accountability. It can be extremely helpful, but it should not approve its own work, merge directly to production, or define business-critical requirements without human oversight.

    A practical adoption path for teams

    The safest way to introduce AI-first development is to start where the cost of being wrong is low and the value of learning is high. Teams do not need to redesign their entire engineering organization on day one.

    • Start with documentation tasks: README updates, setup instructions, changelog drafts, inline comments, and developer onboarding guides.
    • Move to tests: ask agents to generate tests for existing behavior, then have humans review whether the tests reflect reality and cover meaningful cases.
    • Try small bug fixes: choose issues with clear reproduction steps, limited scope, and existing test coverage.
    • Use agents for dependency and compatibility chores: minor version updates, deprecation warnings, formatting changes, and static-analysis cleanup.
    • Experiment with contained refactors: rename internal APIs, simplify duplicate code, or reorganize files where CI can catch regressions.
    • Delay business-critical features: save payments, authentication, permissions, data migrations, and customer-impacting workflows until the team has mature guardrails.

    The first goal is not maximum automation. The first goal is calibration. Teams need to learn which tasks the agent handles well, which prompts produce reliable results, where it fails, and what review checklist catches the most important mistakes.

    Guardrails that make agentic development safer

    Agentic workflows become much more useful when they are surrounded by clear boundaries. The best teams will treat coding agents as part of the software delivery system, not as a side experiment running outside normal controls.

    • Repository instructions: maintain a short, current guide that explains coding style, test commands, architecture rules, naming conventions, and files the agent should not edit without permission.
    • Scoped permissions: limit what the agent can access, execute, or modify. Avoid broad credentials when a read-only or test-only token would work.
    • Branch isolation: require agents to work in separate branches or sandboxed environments instead of editing protected branches directly.
    • Continuous integration checks: run unit tests, linters, type checks, security scans, and build steps before review.
    • Human code review: require a human reviewer for every agent-authored pull request, especially when changes touch security, data, billing, or permissions.
    • Secrets hygiene: prevent agents from reading or printing sensitive keys, customer data, private tokens, or environment files unless there is a specific approved workflow.
    • Evaluation logs: keep records of task prompts, generated diffs, test results, and reviewer feedback so the team can improve prompts and policies over time.
    • Rollback plans: make sure changes can be reverted quickly through version control, feature flags, backups, or deployment controls.

    These controls are not meant to slow everything down. They make it possible to move faster without confusing speed with safety. The more autonomy a tool has, the more important it is to make boundaries explicit.

    A WordPress and plugin-development sidebar

    For CoatiPress readers working in WordPress, coding agents can be especially useful because plugin development often involves repeated patterns: hooks, filters, settings pages, shortcodes, REST routes, admin notices, scripts, styles, sanitization, escaping, and compatibility checks. Those patterns give agents useful context, but they also create security and quality responsibilities that cannot be delegated blindly.

    • Draft tests for plugin functions, REST endpoints, role checks, and settings validation.
    • Review whether hooks and filters are named consistently and documented clearly.
    • Generate documentation for plugin settings, admin screens, and integration steps.
    • Inspect edge cases around logged-in versus logged-out users, API limits, caching, and error handling.
    • Suggest compatibility checks for current WordPress and PHP versions.
    • Flag places where input should be sanitized, output escaped, nonces verified, and capabilities checked.

    For example, an agent might help draft tests for a chat plugin’s logged-in and logged-out request limits, document a content pipeline’s configuration options, or inspect lead-record mapping logic for obvious integration edge cases. But a human developer still owns the release decision, security review, and customer impact.

    The pull request becomes the control point

    AI-first development does not eliminate the pull request. It makes the pull request more important. The PR becomes the place where intent, generated changes, automated checks, risk notes, reviewer comments, and final accountability come together.

    In a mature workflow, the agent should not just dump code. It should explain what it changed, why it changed it, what tests it ran, what it could not verify, and what risks reviewers should inspect. That turns AI output from a mystery patch into a structured engineering artifact.

    What comes next

    Coding agents will keep improving. They will get better at repository context, long-running tasks, test repair, migration planning, and integration with issue trackers and deployment systems. But the winning teams will not be the ones that simply allow the most automation. They will be the ones that design the clearest workflows around it.

    The question for engineering leaders, plugin developers, and technical founders is not whether AI will write code. It already does. The better question is how to turn AI-written code into trustworthy software: scoped tasks, clear context, automated verification, human review, and a culture that treats speed as valuable only when paired with accountability.

    Sources and Fact Check References

    • GitHub Docs – GitHub describes Copilot agent mode as iterating on code, fixing errors, suggesting terminal commands, and analyzing run-time errors.
    • OpenAI – OpenAI positions Codex as a software engineering agent for tasks such as features, bug fixes, refactors, migrations, tests, and code review.
    • Google Jules Docs – Google Jules supports asynchronous coding tasks using connected repositories, branches, generated plans, and reviewable changes.
    • Anthropic – Anthropic provides Claude Code best practices focused on clear context, planning, testing loops, and project-specific instructions.