{"id":19482,"date":"2026-10-01T18:49:49","date_gmt":"2026-10-01T13:19:49","guid":{"rendered":"https:\/\/procreator.design\/blog\/?p=19482"},"modified":"2026-10-01T19:28:58","modified_gmt":"2026-10-01T13:58:58","slug":"agentic-coding-with-claude-codex-skills","status":"publish","type":"post","link":"https:\/\/procreator.design\/blog\/agentic-coding-with-claude-codex-skills\/","title":{"rendered":"How Agentic Coding Uses Claude Code Skills and Codex Skills"},"content":{"rendered":"<p>Most teams think better coding agents need better prompts. The harder problem is deciding what an agent should know, when it should know it, and which parts of the workflow it can execute without a human.<\/p>\n<p>That is where agentic coding is changing software delivery. Claude Code Skills, Codex Skills, repository instructions, MCP tools, scripts, and evals now overlap in the same engineering loop. Without clear boundaries, teams can replace prompt chaos with workflow chaos.<\/p>\n<p>This article explains where each layer belongs, what Claude Code and Codex actually share, which engineering tasks deserve reusable Skills, and how teams can make agent-driven development repeatable without turning every instruction into permanent context.<\/p>\n<h2>TL;DR<\/h2>\n<ul>\n<li>Agentic coding shifts the engineering problem from generating code to designing the context, workflows, tools, checks, and approval boundaries around coding agents.<\/li>\n<li>Repository instructions should hold stable project facts, while Skills should hold repeatable methods that an agent needs only for specific tasks.<\/li>\n<li>Claude Code Skills and Codex Skills support a shared Skill structure, but runtime-specific tools, permissions, and instructions still affect portability.<\/li>\n<li>Skills work best alongside deterministic scripts, MCP tools, evals, and human review rather than replacing every layer with instructions.<\/li>\n<li>Engineering teams should standardize one repetitive, testable workflow before they build a large Skill library.<\/li>\n<\/ul>\n<h2>How Is Agentic Coding Changing Engineering Work?<\/h2>\n<p>Agentic coding moves AI from one step in development into a workflow that can gather context, plan work, edit files, run tools, test results, and return decisions to an engineer.<\/p>\n<p>That is a bigger change than better code completion.<\/p>\n<p>A coding assistant waits for an engineer to decide what happens next. A coding agent can participate in the sequence itself. That makes the surrounding workflow almost as important as the model.<\/p>\n<p><a href=\"https:\/\/www.anthropic.com\/engineering\/harness-design-long-running-apps\" target=\"_blank\" rel=\"noopener\">Anthropic&#8217;s harness research<\/a> shows what changes when the workflow gets more deliberate. Anthropic used a planner, generator, and evaluator for long-running application development. It also broke work into smaller chunks and passed structured artifacts between stages rather than asking one agent to carry the whole build in one undifferentiated task.<\/p>\n<p>The lesson is not that every coding task needs three agents.<\/p>\n<p>Anthropic found that the evaluator still helped on work near the edge of the model&#8217;s capabilities, but became unnecessary overhead on tasks the newer model could already handle reliably.<\/p>\n<p>That creates a different engineering question:<\/p>\n<p><strong>How much structure does this task actually need?<\/strong><\/p>\n<p>Every extra instruction, evaluator, wrapper, or checkpoint creates another assumption the team may eventually need to maintain.<\/p>\n<p>Agentic coding therefore changes two things at once. The model writes more of the code, while the engineering team designs more of the system around how that work happens.<\/p>\n<h2>Why Does Agentic Coding Need More Than AGENTS.md or CLAUDE.md?<\/h2>\n<p><strong>Agentic coding needs more than one permanent instruction file because not every engineering procedure belongs in every task&#8217;s context.<\/strong><\/p>\n<p>Repository instructions still have an important job.<\/p>\n<p>Architecture boundaries, package-management rules, testing commands, naming conventions, and project-wide constraints affect many tasks. They make sense as persistent project context.<\/p>\n<p>A release-verification procedure does not.<\/p>\n<p>Neither does a database migration checklist, dependency-upgrade workflow, or project-specific bug-reproduction process.<\/p>\n<p>The <a href=\"https:\/\/code.claude.com\/docs\/en\/skills\" target=\"_blank\" rel=\"noopener\">Claude Code Skills docs<\/a> make that distinction explicit. Anthropic recommends creating a Skill when engineers repeatedly paste the same checklist or multi-step procedure, or when a section of <code>CLAUDE.md<\/code> has grown into a procedure rather than a fact. Claude loads the Skill body only when the workflow uses it.<\/p>\n<p>That gives teams a cleaner way to organize agent context:<\/p>\n<ul>\n<li><strong>Stable project facts<\/strong> belong in repository instructions because many tasks need them.<\/li>\n<li><strong>Repeatable procedures<\/strong> belong in Skills because only relevant tasks need the full method.<\/li>\n<li><strong>Mechanical checks<\/strong> belong in scripts or hooks because code can verify them consistently.<\/li>\n<li><strong>Current information<\/strong> belongs in tools or MCP connections because a static instruction file will go stale.<\/li>\n<\/ul>\n<p>This is why context engineering should not mean &#8220;write more instructions.&#8221;<\/p>\n<p>It is closer to information architecture for an engineering agent.<\/p>\n<p><img data-recalc-dims=\"1\" loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-large wp-image-19485\" src=\"https:\/\/i0.wp.com\/procreator.design\/blog\/wp-content\/uploads\/2026\/10\/01-Agent-context-architecture.png?resize=1024%2C682&#038;ssl=1\" alt=\"Agent context architecture\" width=\"1024\" height=\"682\" srcset=\"https:\/\/i0.wp.com\/procreator.design\/blog\/wp-content\/uploads\/2026\/10\/01-Agent-context-architecture.png?resize=1024%2C682&amp;ssl=1 1024w, https:\/\/i0.wp.com\/procreator.design\/blog\/wp-content\/uploads\/2026\/10\/01-Agent-context-architecture.png?resize=400%2C266&amp;ssl=1 400w, https:\/\/i0.wp.com\/procreator.design\/blog\/wp-content\/uploads\/2026\/10\/01-Agent-context-architecture.png?resize=768%2C512&amp;ssl=1 768w, https:\/\/i0.wp.com\/procreator.design\/blog\/wp-content\/uploads\/2026\/10\/01-Agent-context-architecture.png?w=1336&amp;ssl=1 1336w\" sizes=\"auto, (max-width: 1000px) 100vw, 1000px\" \/><\/p>\n<p>The distinction also becomes easier once teams separate <a href=\"https:\/\/procreator.design\/blog\/chatgpt-codex-agent-claude-skills-compared\/\" target=\"_blank\" rel=\"noopener\">Claude, Codex, and Agent Skills<\/a> by what belongs to the shared format and what depends on the product running it.<\/p>\n<h2>Where Do Claude Code Skills and Codex Skills Fit?<\/h2>\n<p>Claude Code Skills and Codex Skills sit between persistent project context and the tools that perform work. A Skill gives the agent a reusable method for handling a recognizable task.<\/p>\n<p>Both ecosystems now support the same basic idea: a <code>SKILL.md<\/code> file carries instructions, while supporting files can hold references, scripts, examples, or other resources.<\/p>\n<p>Claude Code states that its Skills follow the open Agent Skills standard. OpenAI likewise says its Skills are compatible with the open Agent Skills standard and can package instructions, references, scripts, and assets.<\/p>\n<p><a href=\"https:\/\/developers.openai.com\/api\/docs\/guides\/tools-skills\" target=\"_blank\" rel=\"noopener\">OpenAI Skills documentation<\/a> also shows why a Skill is different from simply saving a long prompt. OpenAI exposes each available Skill&#8217;s name, description, and path so the agent can discover the relevant workflow, then read its full instructions when needed.<\/p>\n<p>That gives each layer in the engineering stack a different job.<\/p>\n<table style=\"height: 266px;\" border=\"#fff\" width=\"1999\" cellspacing=\"0\">\n<tbody>\n<tr>\n<th style=\"text-align: center;\">Layer<\/th>\n<th style=\"text-align: center;\">What it should carry<\/th>\n<th style=\"text-align: center;\">Engineering example<\/th>\n<\/tr>\n<tr>\n<td style=\"text-align: left; padding: 5px;\"><strong>Task prompt<\/strong><\/td>\n<td style=\"text-align: left; padding: 5px;\">Current objective and acceptance criteria<\/td>\n<td style=\"text-align: left; padding: 5px;\">Fix a checkout retry bug<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left; padding: 5px;\"><strong>Repository instructions<\/strong><\/td>\n<td style=\"text-align: left; padding: 5px;\">Stable project rules<\/td>\n<td style=\"text-align: left; padding: 5px;\">Architecture, test commands, package manager<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left; padding: 5px;\"><strong>Skill<\/strong><\/td>\n<td style=\"text-align: left; padding: 5px;\">Repeatable procedure<\/td>\n<td style=\"text-align: left; padding: 5px;\">PR review, migration review, release verification<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left; padding: 5px;\"><strong>Tool or MCP<\/strong><\/td>\n<td style=\"text-align: left; padding: 5px;\">Current information and controlled actions<\/td>\n<td style=\"text-align: left; padding: 5px;\">GitHub diff, Figma design, issue data<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left; padding: 5px;\"><strong>Script or hook<\/strong><\/td>\n<td style=\"text-align: left; padding: 5px;\">Deterministic validation<\/td>\n<td style=\"text-align: left; padding: 5px;\">Lint, schema validation, generated-file check<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left; padding: 5px;\"><strong>Subagent<\/strong><\/td>\n<td style=\"text-align: left; padding: 5px;\">Isolated or parallel work<\/td>\n<td style=\"text-align: left; padding: 5px;\">Research or independent review<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left; padding: 5px;\"><strong>Eval<\/strong><\/td>\n<td style=\"text-align: left; padding: 5px;\">Evidence that the workflow still behaves correctly<\/td>\n<td style=\"text-align: left; padding: 5px;\">Trigger and outcome tests<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left; padding: 5px;\"><strong>Human gate<\/strong><\/td>\n<td style=\"text-align: left; padding: 5px;\">Consequential judgment<\/td>\n<td style=\"text-align: left; padding: 5px;\">Merge, release, permissions, migration<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p><img data-recalc-dims=\"1\" loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-large wp-image-19486\" src=\"https:\/\/i0.wp.com\/procreator.design\/blog\/wp-content\/uploads\/2026\/10\/02-agenctic-coding-stack.png?resize=1024%2C682&#038;ssl=1\" alt=\"agenctic coding stack\" width=\"1024\" height=\"682\" srcset=\"https:\/\/i0.wp.com\/procreator.design\/blog\/wp-content\/uploads\/2026\/10\/02-agenctic-coding-stack.png?resize=1024%2C682&amp;ssl=1 1024w, https:\/\/i0.wp.com\/procreator.design\/blog\/wp-content\/uploads\/2026\/10\/02-agenctic-coding-stack.png?resize=400%2C266&amp;ssl=1 400w, https:\/\/i0.wp.com\/procreator.design\/blog\/wp-content\/uploads\/2026\/10\/02-agenctic-coding-stack.png?resize=768%2C512&amp;ssl=1 768w, https:\/\/i0.wp.com\/procreator.design\/blog\/wp-content\/uploads\/2026\/10\/02-agenctic-coding-stack.png?w=1336&amp;ssl=1 1336w\" sizes=\"auto, (max-width: 1000px) 100vw, 1000px\" \/><\/p>\n<p>Consider code review.<\/p>\n<p>A repository instruction can tell the agent which test command the project uses. A code-review Skill can define the review sequence, severity model, and reporting format. A script can run linting. GitHub can provide the live diff. An engineer can make the final merge decision.<\/p>\n<p>Each layer has one job.<\/p>\n<p>That is why <a href=\"https:\/\/procreator.design\/blog\/agent-skills-reusable-ai-workflows\/\" target=\"_blank\" rel=\"noopener\">reusable Agent Skills<\/a> make more sense for established procedures than for every piece of information a team wants an agent to remember.<\/p>\n<h2>What Is Actually Portable Between Claude Code Skills and Codex Skills?<\/h2>\n<p><strong>The reusable method is usually more portable than the implementation around it.<\/strong><\/p>\n<p>The shared Skill format makes movement possible. It does not guarantee identical behavior.<\/p>\n<p>A code-review Skill can describe which files to inspect, which risks to check, how to categorize findings, and what evidence the reviewer must provide. That method can make sense in both Claude Code and Codex.<\/p>\n<p>Runtime assumptions move less cleanly.<\/p>\n<p>A Skill may depend on:<\/p>\n<ul>\n<li><strong>Product-specific instructions<\/strong> that only one runtime understands.<\/li>\n<li><strong>Tool names or MCP servers<\/strong> that are not available in the destination environment.<\/li>\n<li><strong>Filesystem paths<\/strong> that assume a particular project or sandbox structure.<\/li>\n<li><strong>Shell commands and packages<\/strong> that need a specific execution environment.<\/li>\n<li><strong>Permission rules<\/strong> that change what the agent can read, execute, or modify.<\/li>\n<\/ul>\n<p>Claude Code, for example, extends the Agent Skills standard with features such as invocation controls, subagent execution, dynamic context injection, and additional frontmatter fields. Anthropic explicitly distinguishes those Claude Code extensions from the fields that belong to the shared specification.<\/p>\n<p>OpenAI&#8217;s implementation has its own execution and discovery model. Its Agents API discovers Skills from registered capability directories inside the sandbox, while uploaded Skills can use versioned bundles.<\/p>\n<p>So portability needs two separate questions:<\/p>\n<ul>\n<li><strong>Can the method move?<\/strong> Often, yes.<\/li>\n<li><strong>Can the implementation move unchanged?<\/strong> Only after the team checks its runtime dependencies.<\/li>\n<\/ul>\n<p><img data-recalc-dims=\"1\" loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-large wp-image-19487\" src=\"https:\/\/i0.wp.com\/procreator.design\/blog\/wp-content\/uploads\/2026\/10\/03-Portable-method-vs-runtime.png?resize=1024%2C682&#038;ssl=1\" alt=\"Portable method vs runtime\" width=\"1024\" height=\"682\" srcset=\"https:\/\/i0.wp.com\/procreator.design\/blog\/wp-content\/uploads\/2026\/10\/03-Portable-method-vs-runtime.png?resize=1024%2C682&amp;ssl=1 1024w, https:\/\/i0.wp.com\/procreator.design\/blog\/wp-content\/uploads\/2026\/10\/03-Portable-method-vs-runtime.png?resize=400%2C266&amp;ssl=1 400w, https:\/\/i0.wp.com\/procreator.design\/blog\/wp-content\/uploads\/2026\/10\/03-Portable-method-vs-runtime.png?resize=768%2C512&amp;ssl=1 768w, https:\/\/i0.wp.com\/procreator.design\/blog\/wp-content\/uploads\/2026\/10\/03-Portable-method-vs-runtime.png?w=1336&amp;ssl=1 1336w\" sizes=\"auto, (max-width: 1000px) 100vw, 1000px\" \/><\/p>\n<p>The first question is about knowledge. The second is engineering.<\/p>\n\n<h2>What Does a Mature Agentic Coding Workflow Look Like?<\/h2>\n<p><strong>A mature agentic coding workflow moves from a defined outcome to verified evidence. It does not stop when the agent produces code.<\/strong><\/p>\n<p>A practical workflow can follow eight steps:<\/p>\n<ol start=\"1\">\n<li><strong>Define the outcome.<\/strong> Give the agent concrete acceptance criteria instead of asking it to &#8220;improve&#8221; an implementation.<\/li>\n<li><strong>Load project constraints.<\/strong> Repository instructions establish architecture, conventions, and non-negotiable rules.<\/li>\n<li><strong>Select the relevant method.<\/strong> The Skill supplies the procedure for this type of task.<\/li>\n<li><strong>Retrieve current evidence.<\/strong> Tools provide the diff, ticket, Figma file, logs, or environment state.<\/li>\n<li><strong>Execute bounded work.<\/strong> The agent edits code or delegates isolated tasks without expanding the scope unnecessarily.<\/li>\n<li><strong>Run deterministic checks.<\/strong> Tests, linters, type checks, and validators catch failures that do not need model judgment.<\/li>\n<li><strong>Evaluate the result.<\/strong> The workflow compares actual behavior with the original acceptance criteria.<\/li>\n<li><strong>Escalate consequential choices.<\/strong> Engineers retain control over production releases, schema changes, architectural exceptions, and other high-impact decisions.<\/li>\n<\/ol>\n<p>The ordering matters.<\/p>\n<p>If a linter can answer a question deterministically, asking a model to &#8220;check carefully&#8221; creates unnecessary uncertainty.<\/p>\n<p>If a connected system can provide current production or design context, copying an old snapshot into a Skill creates unnecessary staleness.<\/p>\n<p>If an engineer must approve a database migration anyway, pretending the entire workflow is autonomous adds little value.<\/p>\n<p>The goal is not maximum autonomy.<\/p>\n<p>The goal is a clear division of responsibility.<\/p>\n<p><img data-recalc-dims=\"1\" loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-large wp-image-19488\" src=\"https:\/\/i0.wp.com\/procreator.design\/blog\/wp-content\/uploads\/2026\/10\/04-Agentic-coding-workflow.png?resize=1024%2C682&#038;ssl=1\" alt=\"Agentic coding workflow\" width=\"1024\" height=\"682\" srcset=\"https:\/\/i0.wp.com\/procreator.design\/blog\/wp-content\/uploads\/2026\/10\/04-Agentic-coding-workflow.png?resize=1024%2C682&amp;ssl=1 1024w, https:\/\/i0.wp.com\/procreator.design\/blog\/wp-content\/uploads\/2026\/10\/04-Agentic-coding-workflow.png?resize=400%2C266&amp;ssl=1 400w, https:\/\/i0.wp.com\/procreator.design\/blog\/wp-content\/uploads\/2026\/10\/04-Agentic-coding-workflow.png?resize=768%2C512&amp;ssl=1 768w, https:\/\/i0.wp.com\/procreator.design\/blog\/wp-content\/uploads\/2026\/10\/04-Agentic-coding-workflow.png?w=1336&amp;ssl=1 1336w\" sizes=\"auto, (max-width: 1000px) 100vw, 1000px\" \/><\/p>\n<p>Teams designing that division can use ProCreator&#8217;s <a href=\"https:\/\/procreator.design\/services\/ai-innovation-strategy\" target=\"_blank\" rel=\"noopener\">AI Innovation Strategy<\/a> to map what the model should handle, what code should validate, which systems should provide live context, and where human approval should remain explicit.<\/p>\n<h2>Which Engineering Workflows Should Become Skills First?<\/h2>\n<p><strong>The best first Skills are repetitive workflows where the engineering team already understands the method and can inspect the result objectively.<\/strong><\/p>\n<p>Strong candidates include:<\/p>\n<ul>\n<li><strong>Pull request review:<\/strong> Capture repository-specific checks, recurring failure patterns, and evidence expectations.<\/li>\n<li><strong>Database migration review:<\/strong> Standardize sequencing, rollback requirements, schema checks, and approval boundaries.<\/li>\n<li><strong>Bug reproduction:<\/strong> Define environment setup, evidence collection, reproduction steps, and verification.<\/li>\n<li><strong>Release verification:<\/strong> Confirm that the running application behaves correctly instead of treating a passing test suite as the finish line.<\/li>\n<li><strong>Dependency upgrades:<\/strong> Apply project-specific compatibility checks and regression tests.<\/li>\n<li><strong>Design system implementation:<\/strong> Guide agents toward approved components, tokens, and contribution rules.<\/li>\n<li><strong>Accessibility review:<\/strong> Combine automated checks with the parts that still need engineering judgment.<\/li>\n<\/ul>\n<p>Coinbase provides a useful example of the design-system case.<\/p>\n<p>According to <a href=\"https:\/\/www.figma.com\/blog\/how-coinbase-used-code-connect-to-shrink-token-costs\/\" target=\"_blank\" rel=\"noopener\">Figma&#8217;s Coinbase case study<\/a>, the Coinbase Design System team already used agent Skills to teach coding agents which components to select, which deprecated patterns to avoid, and which design tokens to use. The team then added Code Connect so agents could start with production component context rather than searching and guessing.<\/p>\n<p>Figma measured the difference across three controlled runs. Code Connect reduced token use by <strong>11.5 percent<\/strong>, implementation time by <strong>22.3 percent<\/strong>, and cost by <strong>22.5 percent<\/strong> on average. Figma also reported stronger component selection and design-system adherence.<\/p>\n<p>That example separates two jobs teams often combine:<\/p>\n<ul>\n<li><strong>Agent Skills shape the method.<\/strong> They tell the agent how the team expects the implementation to work.<\/li>\n<li><strong>Code Connect supplies context.<\/strong> It gives the agent production component information before it starts making implementation decisions.<\/li>\n<\/ul>\n<p>Neither replaces the other.<\/p>\n<p><img data-recalc-dims=\"1\" loading=\"lazy\" decoding=\"async\" class=\"aligncenter size-large wp-image-19489\" src=\"https:\/\/i0.wp.com\/procreator.design\/blog\/wp-content\/uploads\/2026\/10\/05-method-vs-context.png?resize=1024%2C682&#038;ssl=1\" alt=\"method vs context\" width=\"1024\" height=\"682\" srcset=\"https:\/\/i0.wp.com\/procreator.design\/blog\/wp-content\/uploads\/2026\/10\/05-method-vs-context.png?resize=1024%2C682&amp;ssl=1 1024w, https:\/\/i0.wp.com\/procreator.design\/blog\/wp-content\/uploads\/2026\/10\/05-method-vs-context.png?resize=400%2C266&amp;ssl=1 400w, https:\/\/i0.wp.com\/procreator.design\/blog\/wp-content\/uploads\/2026\/10\/05-method-vs-context.png?resize=768%2C512&amp;ssl=1 768w, https:\/\/i0.wp.com\/procreator.design\/blog\/wp-content\/uploads\/2026\/10\/05-method-vs-context.png?w=1336&amp;ssl=1 1336w\" sizes=\"auto, (max-width: 1000px) 100vw, 1000px\" \/><\/p>\n<p>The same distinction matters when teams connect <a href=\"https:\/\/procreator.design\/blog\/how-to-use-figma-mcp\/\" target=\"_blank\" rel=\"noopener\">Figma MCP with design systems<\/a>. Live design information should come from the connected source. The reusable engineering procedure should sit around that context.<\/p>\n<p>A good Skill does not contain everything.<\/p>\n<p>It contains the method worth repeating.<\/p>\n<h2>How Should Teams Test and Govern Skills?<\/h2>\n<p><strong>Teams should test Skills as engineering behavior, not approve them only because the instructions read well.<\/strong><\/p>\n<p>A reasonable <code>SKILL.md<\/code> can still trigger on the wrong request, skip a required action, or create something outside the expected scope.<\/p>\n<p>A basic Skill test should cover:<\/p>\n<table style=\"height: 266px;\" border=\"#fff\" width=\"1999\" cellspacing=\"0\">\n<tbody>\n<tr>\n<th style=\"text-align: center;\">Test<\/th>\n<th style=\"text-align: center;\">What to check<\/th>\n<\/tr>\n<tr>\n<td style=\"text-align: left; padding: 5px;\"><strong>Trigger precision<\/strong><\/td>\n<td style=\"text-align: left; padding: 5px;\">Does the Skill activate for the intended task?<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left; padding: 5px;\"><strong>False positives<\/strong><\/td>\n<td style=\"text-align: left; padding: 5px;\">Does it stay inactive for unrelated work?<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left; padding: 5px;\"><strong>Required process<\/strong><\/td>\n<td style=\"text-align: left; padding: 5px;\">Did the agent perform the expected checks?<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left; padding: 5px;\"><strong>Outcome<\/strong><\/td>\n<td style=\"text-align: left; padding: 5px;\">Did the result meet the definition of done?<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left; padding: 5px;\"><strong>Side effects<\/strong><\/td>\n<td style=\"text-align: left; padding: 5px;\">Did the agent modify anything outside scope?<\/td>\n<\/tr>\n<tr>\n<td style=\"text-align: left; padding: 5px;\"><strong>Efficiency<\/strong><\/td>\n<td style=\"text-align: left; padding: 5px;\">Did the workflow add instructions or actions that no longer improve the result?<\/td>\n<\/tr>\n<\/tbody>\n<\/table>\n<p>Coinbase&#8217;s design-system work shows the value of measuring this instead of trusting intuition. The team ran controlled comparisons with the same design, prompt, and model, then measured token use, implementation time, cost, component choice, and design-system adherence. Figma also reports that the team regularly evaluates its agent Skills as the models change.<\/p>\n<p>That last part matters.<\/p>\n<p>Skills encode assumptions about what an agent needs help with. Model capabilities change. Tooling changes. Repository conventions change.<\/p>\n<p>So Skill maintenance should sometimes remove instructions rather than add them.<\/p>\n<p>Teams should:<\/p>\n<ul>\n<li><strong>Delete stale rules<\/strong> when the model or workflow no longer needs them.<\/li>\n<li><strong>Split overlapping Skills<\/strong> when agents struggle to choose the right procedure.<\/li>\n<li><strong>Move mechanical checks into scripts<\/strong> when code can produce a clear pass or fail.<\/li>\n<li><strong>Retest important workflows<\/strong> after model, tool, or permission changes.<\/li>\n<li><strong>Assign an owner<\/strong> when a Skill influences consequential engineering work.<\/li>\n<\/ul>\n<p>The library should get better as it grows, not simply bigger.<\/p>\n<h2>What Should Teams Standardize Next for Agentic Coding?<\/h2>\n<p><strong>Teams adopting agentic coding should standardize one bounded workflow before they build an entire library of Skills.<\/strong><\/p>\n<p>Pick something frequent enough to matter and clear enough to evaluate. Code review, release verification, migration review, dependency upgrades, and design-system implementation all give the team a better starting point than a generic &#8220;software engineering Skill.&#8221;<\/p>\n<p>Then watch what the agent repeatedly needs, what code can verify without interpretation, and where engineers still need to make the call. Those boundaries will tell you what should become reusable infrastructure.<\/p>\n<p>The future of agentic coding will not belong to the team with the longest instruction file. It will belong to teams that know what context matters, when it matters, and when the agent should stop.<\/p>\n<p>If your team already has an AI-assisted engineering workflow that needs clearer context, validation, or human decision boundaries, <a href=\"https:\/\/procreator.design\/contact-us\/start-project-primary\" target=\"_blank\" rel=\"noopener\">start a project with ProCreator<\/a>.<\/p>\n\n<h3>FAQs<\/h3>\n<style>#sp-ea-19493 .spcollapsing { height: 0; overflow: hidden; transition-property: height;transition-duration: 300ms;}#sp-ea-19493.sp-easy-accordion>.sp-ea-single {margin-bottom: 10px; border: 1px solid #e2e2e2; }#sp-ea-19493.sp-easy-accordion>.sp-ea-single>.ea-header a {color: #444;}#sp-ea-19493.sp-easy-accordion>.sp-ea-single>.sp-collapse>.ea-body {background: #fff; color: #444;}#sp-ea-19493.sp-easy-accordion>.sp-ea-single {background: #eee;}#sp-ea-19493.sp-easy-accordion>.sp-ea-single>.ea-header a .ea-expand-icon { float: left; color: #444;font-size: 16px;}<\/style><div id=\"sp_easy_accordion-1790861463\"><div id=\"sp-ea-19493\" class=\"sp-ea-one sp-easy-accordion\" data-ea-active=\"ea-click\" data-ea-mode=\"vertical\" data-preloader=\"\" data-scroll-active-item=\"\" data-offset-to-scroll=\"0\"><div class=\"ea-card ea-expand sp-ea-single\"><h3 class=\"ea-header\"><a class=\"collapsed\" id=\"ea-header-194930\" role=\"button\" data-sptoggle=\"spcollapse\" data-sptarget=\"#collapse194930\" aria-controls=\"collapse194930\" href=\"#\" aria-expanded=\"true\" tabindex=\"0\"><i aria-hidden=\"true\" role=\"presentation\" class=\"ea-expand-icon eap-icon-ea-expand-minus\"><\/i> What is agentic coding?<\/a><\/h3><div class=\"sp-collapse spcollapse collapsed show\" id=\"collapse194930\" data-parent=\"#sp-ea-19493\" role=\"region\" aria-labelledby=\"ea-header-194930\"> <div class=\"ea-body\"><p>Agentic coding is a software-development approach where AI agents participate in multi-step engineering work instead of only suggesting code. An agent may gather repository context, plan changes, edit files, use tools, run checks, and verify results. Engineers still define its authority, acceptance criteria, and approval boundaries.Agentic coding is a software-development approach where AI agents participate in multi-step engineering work instead of only suggesting code. An agent may gather repository context, plan changes, edit files, use tools, run checks, and verify results. Engineers still define its authority, acceptance criteria, and approval boundaries.<\/p><\/div><\/div><\/div><div class=\"ea-card sp-ea-single\"><h3 class=\"ea-header\"><a class=\"collapsed\" id=\"ea-header-194931\" role=\"button\" data-sptoggle=\"spcollapse\" data-sptarget=\"#collapse194931\" aria-controls=\"collapse194931\" href=\"#\" aria-expanded=\"false\" tabindex=\"0\"><i aria-hidden=\"true\" role=\"presentation\" class=\"ea-expand-icon eap-icon-ea-expand-plus\"><\/i> What is the difference between Claude Code Skills and Codex Skills?<\/a><\/h3><div class=\"sp-collapse spcollapse \" id=\"collapse194931\" data-parent=\"#sp-ea-19493\" role=\"region\" aria-labelledby=\"ea-header-194931\"> <div class=\"ea-body\"><p>Claude Code Skills and Codex Skills both package reusable methods around engineering tasks, and both support the open Agent Skills approach. Their runtimes differ. Tool access, permissions, filesystem behavior, invocation rules, and platform-specific features can change how the same procedure behaves after a team moves it between environments.<\/p><\/div><\/div><\/div><div class=\"ea-card sp-ea-single\"><h3 class=\"ea-header\"><a class=\"collapsed\" id=\"ea-header-194932\" role=\"button\" data-sptoggle=\"spcollapse\" data-sptarget=\"#collapse194932\" aria-controls=\"collapse194932\" href=\"#\" aria-expanded=\"false\" tabindex=\"0\"><i aria-hidden=\"true\" role=\"presentation\" class=\"ea-expand-icon eap-icon-ea-expand-plus\"><\/i> Should every engineering workflow become a Skill?<\/a><\/h3><div class=\"sp-collapse spcollapse \" id=\"collapse194932\" data-parent=\"#sp-ea-19493\" role=\"region\" aria-labelledby=\"ea-header-194932\"> <div class=\"ea-body\"><p>No. A task should become a Skill when it repeats, follows a recognizable method, needs task-specific context, and produces something the team can verify. Stable project facts belong in repository instructions. Deterministic checks belong in scripts. One-off work usually belongs in the current task prompt.<\/p><p>&nbsp;<\/p><\/div><\/div><\/div><div class=\"ea-card sp-ea-single\"><h3 class=\"ea-header\"><a class=\"collapsed\" id=\"ea-header-194933\" role=\"button\" data-sptoggle=\"spcollapse\" data-sptarget=\"#collapse194933\" aria-controls=\"collapse194933\" href=\"#\" aria-expanded=\"false\" tabindex=\"0\"><i aria-hidden=\"true\" role=\"presentation\" class=\"ea-expand-icon eap-icon-ea-expand-plus\"><\/i> How many Skills should an engineering repository have?<\/a><\/h3><div class=\"sp-collapse spcollapse \" id=\"collapse194933\" data-parent=\"#sp-ea-19493\" role=\"region\" aria-labelledby=\"ea-header-194933\"> <div class=\"ea-body\"><p>There is no useful universal number. Keep the Skills that improve recurring work enough to justify their context and maintenance cost. A smaller set with distinct triggers usually gives the agent clearer choices than a catalog full of overlapping procedures that solve nearly the same problem.<\/p><p>&nbsp;<\/p><\/div><\/div><\/div><div class=\"ea-card sp-ea-single\"><h3 class=\"ea-header\"><a class=\"collapsed\" id=\"ea-header-194934\" role=\"button\" data-sptoggle=\"spcollapse\" data-sptarget=\"#collapse194934\" aria-controls=\"collapse194934\" href=\"#\" aria-expanded=\"false\" tabindex=\"0\"><i aria-hidden=\"true\" role=\"presentation\" class=\"ea-expand-icon eap-icon-ea-expand-plus\"><\/i> How do you test a Claude Code or Codex Skill?<\/a><\/h3><div class=\"sp-collapse spcollapse \" id=\"collapse194934\" data-parent=\"#sp-ea-19493\" role=\"region\" aria-labelledby=\"ea-header-194934\"> <div class=\"ea-body\"><p>Test whether the Skill activates for the right tasks, stays inactive for adjacent requests, follows the required procedure, produces the expected result, and avoids unwanted side effects. Keep representative regression tasks and rerun them when the Skill, its connected tools, permissions, or underlying model changes.<\/p><p>&nbsp;<\/p><\/div><\/div><\/div><\/div><\/div>\n","protected":false},"excerpt":{"rendered":"<p>Most teams think better coding agents need better prompts. The harder problem is deciding what an agent should know, when it should know it, and which parts of the workflow it can execute without a human. That is where agentic coding is changing software delivery. Claude Code Skills, Codex Skills, repository instructions, MCP tools, scripts, [&hellip;]<\/p>\n","protected":false},"author":6,"featured_media":19491,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[63],"tags":[],"class_list":["post-19482","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ai-and-innovation"],"_links":{"self":[{"href":"https:\/\/procreator.design\/blog\/wp-json\/wp\/v2\/posts\/19482","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/procreator.design\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/procreator.design\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/procreator.design\/blog\/wp-json\/wp\/v2\/users\/6"}],"replies":[{"embeddable":true,"href":"https:\/\/procreator.design\/blog\/wp-json\/wp\/v2\/comments?post=19482"}],"version-history":[{"count":4,"href":"https:\/\/procreator.design\/blog\/wp-json\/wp\/v2\/posts\/19482\/revisions"}],"predecessor-version":[{"id":19494,"href":"https:\/\/procreator.design\/blog\/wp-json\/wp\/v2\/posts\/19482\/revisions\/19494"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/procreator.design\/blog\/wp-json\/wp\/v2\/media\/19491"}],"wp:attachment":[{"href":"https:\/\/procreator.design\/blog\/wp-json\/wp\/v2\/media?parent=19482"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/procreator.design\/blog\/wp-json\/wp\/v2\/categories?post=19482"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/procreator.design\/blog\/wp-json\/wp\/v2\/tags?post=19482"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}