2026 Agentic Software Engineering: A Terminology and Literature Placeholder
Table of Contents
1. Overview
This is a placeholder. The terminology below is unsettled as of September 2026: different papers and vendors name overlapping pieces of the same problem differently, and no umbrella term has stabilized yet.
The problem itself is concrete: many AI-agent-authored pull requests compete to land on one trunk. The umbrella term that appears most often for the broader practice is agentic software engineering, sometimes called "SE 3.0." The specific problem — many agent-authored changes competing for one trunk — is usually called integration of agent-authored pull requests, or the integration bottleneck.
2. Names by sub-problem
No single name covers all five parts. Each sub-problem already has an established name from pre-agent software engineering research, and a newer, narrower name that 2026 papers are using for the agent-specific case.
| Sub-problem | Established name | 2026 agent-specific name |
|---|---|---|
| Ordering and batching merges | Merge queues; "keeping master green"; speculative CI | Agent-scale or speculative merge queues |
| Concurrent changes that conflict | Merge conflict prediction; semi-structured merge | Cross-agent conflicts; merge contention |
| Predicting that a change will break production | Just-in-time (JIT) defect prediction; change risk assessment | Merge-outcome predictors for agent PRs |
| Deciding whether a change may deploy | Release engineering; change management (ITIL change enablement) | Agent governance; deploy gating |
| Coordinating agents before they write code | Task decomposition; concurrency control | Pre-write admission; multi-agent code co-synthesis |
3. 2026 research
3.1. Conflict measurement
AgenticFlict (Ogenrwot and Businge 2026) (Ogenrwot and Businge, arXiv:2604.03551, AIware 2026) simulates merging each agent PR against its base branch. Of more than 142,000 agent PRs collected across more than 59,000 repositories, more than 107,000 were successfully processed through the merge simulation; 27.67% of those processed PRs conflicted. Conflict rate increases with PR size (lines added plus deleted), and rates differ by agent.
AI Agent Pull Requests on GitHub: Frequency, Structure, and Merge Conflict Rates (Xu, Subramanian, and Karthik 2026) (arXiv:2607.04697) uses the AIDev-pop dataset. In 40.2% of repositories, some pairs of agent PRs were open at the same time, and those pairs account for 79.4% of all agent PRs. The authors list lost CI compute and token spend on conflict resolution as hypothesized costs; they did not measure them.
3.2. Merge-outcome prediction
When AI Teammates Meet Code Review (Nachuma and Zibran 2026) (Nachuma and Zibran, MSR 2026, arXiv:2602.19441) fits a logistic regression to 33,596 agent-authored PRs across five coding agents. 71.5% of these PRs merged overall, but the share varies sharply by agent: OpenAI Codex 82.6%, Devin 53.8%, Copilot 43.0%. Receiving at least one human review before the merge decision is the strongest positive predictor of merge.
3.3. Coordination before writing
ATM: CID-Brokered Pre-Write Admission for Multi-Agent Code Co-Synthesis (Huang 2026) (arXiv:2607.00041) admits agent writes to shared code before they happen. It applies concurrency-control ideas (read-set tracking, isolation levels) to agents writing code, rather than resolving conflicts at merge time.
3.4. Trust boundaries
Knowledge-Based Pull Requests (Zhang and Sun 2026) (arXiv:2606.26721) proposes that an external contributor's code and agent trace serve as input knowledge. A project-owned agent then regenerates the code inside the receiving project. This separates two decisions: whether the knowledge should enter the project, and whether a specific implementation should merge.
3.5. Merge queues at agent scale
Tian Pan's July 2026 essay applies queueing theory to a serial merge queue. With 30-minute CI, the queue lands at most two PRs per hour. Agents raise the arrival rate while the service rate stays fixed.
The same essay notes that batch failures force bisection, which costs O(log n) CI runs. Flaky tests set how often this happens, so agents that re-queue failed PRs increase the cost.
Mergify's speculative checks test the cumulative merges (PR1), (PR1+PR2), (PR1+PR2+PR3) in parallel lanes.
4. Pre-agent foundations
The agent-specific work above builds on an older, non-agent-specific literature:
- Ananthanarayanan et al., "Keeping Master Green at Scale" (Ananthanarayanan et al. 2019) (EuroSys 2019), describes Uber's SubmitQueue. SubmitQueue predicts build outcomes to schedule speculative builds; it is the main academic treatment of merge scheduling.
- Kamei et al., "A Large-Scale Empirical Study of Just-in-Time Quality Assurance" (Kamei et al. 2013) (TSE 2013), is the standard reference for JIT defect prediction: assigning a risk score to each commit.
- Google's TAP publications cover presubmit/postsubmit testing and culprit finding at monorepo scale.
- Ghiotto et al., "On the Nature of Merge Conflicts" (Ghiotto et al. 2020) (TSE 2020), and Apel et al.'s semistructured merge (Apel et al. 2011) (FSE 2011) cover conflict characterization and resolution.
- The DORA metrics (change failure rate, lead time) and ITIL 4 change enablement (standard, normal, and emergency changes; the change schedule) are the operations-side terms for risk-gated integration.
5. Search terms
A practical list for finding more of this literature: "agent-authored pull requests", "AIDev dataset", "agentic PR integration", "merge queue speculative", "keeping master green", "just-in-time defect prediction", "change risk assessment", "multi-agent code co-synthesis", and "SE 3.0". MSR 2026 and AIware 2026 are the venues where the agent-specific work is concentrated.
6. Closing synthesis
Two other projects were mentioned in passing as covering pieces of this problem: a "slipway" deploy-gating pipeline for the change-risk and ITIL-change-enablement part, and a "wake" post-merge survival analysis for the outcome-measurement part. Neither name matches anything findable in this repository's own site content or history (checked via search and a full-tree grep) — they read as external projects, not something documented here, so they are reported as context rather than linked.
7. References
Seen at
- Atlassian Team '26 Europe: Closing the context gap — a risk assessment surfaced in Jira to fast-track low-risk agent changes.
- Atlassian Team '26 Europe: Bitbucket for the AI era — Bitbucket agents and the delivery platform for agent-authored change.
- Boston AI Week 2026: The Multi-Agent Symphony — review, testing and merge strategies when several coding agents write to one codebase at once.