Encoding your judgment into Claude Code Skills and Hooks

How Much of a Developer's Work Can Be Automated?

Putting my judgment criteria into Claude Code’s Skills and Hooks, and automating starting from small tasks

AI Writes the Code, So Why Am I Still Busy?

When you ask AI to write code, you get results pretty quickly.

But looking back at the whole development process, there are still many things a human has to take care of.

I copy and paste error logs, point out related files, and explain the project rules again. When code comes out, I ask it to run the tests, and if they fail, I paste the results back. When the work is done, I summarize the changes and update the docs.

“Check this too.”

“Run the tests too.”

“Match the existing code style.”

I’m clearly using AI, yet at some point I find myself dictating AI’s every next move. Thinking about this, a question naturally arises.

Can’t we automate even the process of giving repeated instructions?

To me, the core of automation is getting the AI agent to engineer based on the same thought process as mine.

When facing a problem, what do I check first, on what basis do I decide the fix direction, and how far must I verify before calling it done? I want to put these criteria I use while developing into the agent’s workflow as well. Whether those criteria are actually applied can be examined through the evidence the agent checked, the changes it made, and its verification results.

Of course, it’s hard to get this right in one go. So I think a good approach is to automate one small task first, verify that it was performed as intended, and then expand what worked to the next task.

In a previous post, I talked about context and development flow for using AI coding tools. This time, I’ll use Claude Code’s Skills and Hooks to turn those judgment criteria into work procedures. The flow starts with fixing errors, then applies the verified approach to new feature development.

The examples assume a Java, Spring Boot, and Gradle project. It’s aimed at developers with basic Claude Code experience, and introduces a setup you can adapt to your own project.

Verification scope of the examples: I checked the settings JSON, Bash syntax, and the success/failure behavior of the check script in a temporary Git repository. This is not a case of running a full bug fix end-to-end in an actual Claude Code session on a Spring Boot project.

1. First, Slice Off the Task to Automate

Starting with “I’ll automate all of development” makes the scope far too wide. It’s better to first pick one small task where I can explain the order of work and the reasoning behind my decisions. For example, a single bug with clear reproduction steps, or a single check I always run after a fix.

Starting small also makes the results easier to examine. You can see where the agent judged differently from what you intended, and what information or criteria you need to provide more of. The tasks below are also better seen as targets to apply and verify one by one, rather than connecting them all at once.

Task Required input Result to check
Bug fix Reproduction steps, error logs, expected behavior Reproduction test and post-fix verification results
Add an API Request/response spec, domain rules Implementation, tests, API docs
Code review Change diff, requirements Issues with evidence and improvement suggestions
Write PR description Final changes, verification results Purpose, changes, and verification method

For example, “fix the bug” has no definition of done. On the other hand, “reproduce the issue where an invalid sort value causes a 500, and fix it to return 400” has a behavior to check.

This is where automation begins: defining together what the agent should do, the criteria it should use to judge, and the evidence that the work is done.

2. Stop Explaining the Project Every Time

To hand work over to an agent, it first needs to understand the working environment.

What the build command is, where exceptions are handled, whether tests need a DB. It’s not much different from what a person checks when joining a new project.

In Claude Code, you can write these project instructions in CLAUDE.md. Rules for specific concerns can be split into .claude/rules/. However, instructions written in a document alone don’t guarantee compliance, so it’s better to pair verifiable rules with checking tools. [1]

2.1. Put Common Context in CLAUDE.md

Here’s a simple example of a CLAUDE.md placed at the project root.

# Project Guide

## Tech and Structure
- Uses Java 21, Spring Boot, and Gradle.
- Follow the existing Controller → Service → Repository structure.
- Exception responses are handled in GlobalExceptionHandler.

## Running and Verification
- All tests: ./gradlew test
- Single test: ./gradlew test --tests 'package.TestName'
- See docs/testing.md for the environment needed to run tests.

## Work Rules
- Check similar existing code and tests before implementing.
- Propose refactoring unrelated to the requirements as a separate task.
- If you change behavior, verify that behavior.
- Note the reason for any verification you couldn't run in the completion report.

2.2. Leave Execution Conditions and the Reasoning Behind Decisions

You don’t need to put every design document and past discussion here. Keep frequently needed info short, and have it reference related docs for details.

It’s especially troublesome to list only the test command and leave out the conditions needed to run it. The agent needs to know about the test DB, environment variables, and whether external APIs are mocked to distinguish code errors from environment problems.

When writing rules, it’s also worth reflecting on why I work the way I do. A rule like “check similar existing code first” carries the intent of understanding the project’s design and making changes that fit the existing structure. Leaving enough of this reasoning makes it easier for the agent to apply the criteria in new situations.

3. What Roles Do Skills and Hooks Play?

By name alone, both sound like automation features. In an actual setup, it’s easier to understand them by splitting their roles.

Component Role Example
CLAUDE.md / Rules Common project context and rules Exception response conventions, test commands
Skills Work procedures to repeat Bug investigation → fix → verification order
Hooks Actions tied to specific events Run a check after a file edit
CLI / MCP Access to tools and data needed for the work Running tests, fetching issues and logs

3.1. Skills: Procedures and Judgment Criteria to Repeat

A Skill is a feature where you write instructions in SKILL.md. You can configure it so Claude loads it for related tasks, or the user can call it directly with /skill-name. Since it can provide a detailed procedure at the right moment, it saves you from pasting long prompts every time. [2]

3.2. Hooks: When Checks Run

A Hook is an action tied to a specific Claude Code event. In this post, I’ll use a command Hook that runs a shell script when an event occurs. [3]

For example, you can split the roles like this:

Put “the order of investigating and fixing a failure” in a Skill, and “when to run checks after a file edit” in a Hook.

Interpreting check results and deciding the next fix is the agent’s role. Setting up a Hook doesn’t guarantee that the bug fix itself is correct.

4. Turning the Error-Fixing Procedure into a Skill

4.1. Write Down How I Solve Problems

If you start with fixing errors, you can first write down: “What would I check if I took on this problem myself?” I’d check the expected behavior, read the logs, and try to reproduce the part I suspect is the cause. After the fix, I’d check that the same problem is gone and that surrounding behavior isn’t affected.

Let’s move this sequence of judgments into a Skill. For example, the reason to reproduce before fixing is to confirm that the suspected cause actually connects to the real problem. When each step has a purpose like this, it’s also clear what to look at when reviewing the results.

The order for this example is:

  1. Compare expected behavior with actual behavior.
  2. Check logs and related code.
  3. If possible, reproduce the problem with a failing test.
  4. Fix the part responsible for the cause.
  5. Run the reproduction test and related tests.
  6. Report the changes and remaining uncertainties.

4.2. Writing the Bug Fix Skill

Write the following in the project’s .claude/skills/fix-bug/SKILL.md.

---
name: fix-bug
description: Use when investigating and fixing a bug that has reproduction steps or failure logs.
disable-model-invocation: true
---

# Bug Fix

Request: $ARGUMENTS

## Procedure
1. Read CLAUDE.md and related rules.
2. Summarize expected behavior and actual behavior.
3. Check related code, tests, and provided logs.
4. If possible, write a failing reproduction test before fixing.
5. Modify code within the scope needed to resolve the confirmed cause.
6. Run the reproduction test and related affected tests.
7. Check the final diff for unnecessary changes.

## Judgment Criteria
- Separate facts confirmed from logs and cause hypotheses not yet confirmed.
- Base the fix direction on reproduction results or evidence from related code.
- Tests must verify expected behavior and must fail on the wrong behavior.
- If results differ from expectations, re-examine the cause hypothesis.

## When to Stop and Ask
- If expected behavior is ambiguous, ask before implementing.
- If fixes for the same failure fail 3 times in a row,
  report what was tried and what additional information is needed.
- Distinguish external service or test environment problems from code defects.
- Do not delete tests or arbitrarily lower expected values to make them pass.

## Completion Report
- Cause and the evidence for that judgment
- Changes
- Verifications run and their results
- Verifications not run and why

disable-model-invocation: true is a setting that makes the Skill invocable only explicitly by the user. In this example, I added it to intentionally start the bug fix procedure. [2]

4.3. Apply It to a Small Bug and Check the Results

You can invoke it like this:

/fix-bug Passing an unsupported sort value to the product list API
causes a 500. Fix it to return 400 following the project's existing
exception response format, and add a regression test.

This request includes the existing response format and a test to prevent recurrence in the scope of work. I passed along the criteria I would care about if I were fixing it myself.

On the first run, review the reasoning along with the final code. Check whether it actually reproduced the problem, whether it fixed the part matching the cause, and whether the added test can catch the original error.

The retry count written in the Skill is an instruction for the model to follow. If you need a strict cap on executions or cost, you must also enforce it in separate execution control code.

5. Connecting Check Results to the Work with Hooks

Even with a work procedure in place, small checks can be repeatedly missed. Let’s connect those checks to a Hook first.

Here, I’ll use an example that checks Git-tracked changes for whitespace errors and the like after Claude edits a file with the Write or Edit tool. It’s lighter than a functional test and simple enough to confirm how Hooks work.

5.1. Attach a Check to the File Edit Event

Add the following to the project’s .claude/settings.json. If you already have settings, merge the hooks section.

{
  "hooks": {
    "PostToolUse": [
      {
        "matcher": "Write|Edit",
        "hooks": [
          {
            "type": "command",
            "command": "bash \"$CLAUDE_PROJECT_DIR/.claude/hooks/check-diff.sh\"",
            "timeout": 10
          }
        ]
      }
    ]
  }
}

5.2. Pass Check Failures to the Agent

The .claude/hooks/check-diff.sh to connect is as follows. It assumes an environment with Bash and Git.

#!/usr/bin/env bash
set -u

cd "${CLAUDE_PROJECT_DIR:?}" || exit 1

# This example doesn't use the event input.
cat >/dev/null

check_output=$(git diff --check 2>&1)
check_status=$?

if [ "$check_status" -eq 0 ]; then
  exit 0
fi

printf '%s\n' "$check_output" >&2
printf '%s\n' \
  'Change check failed. Please check the output for the cause.' >&2
exit 2

Returning exit code 2 from PostToolUse passes the stderr content to Claude. It doesn’t undo the file edit that already happened. It connects the feedback so the agent can decide its next action. [3]

5.3. Adjust Check Scope and Execution Cost

This example has limits. git diff --check by default checks unstaged changes in tracked files. It doesn’t check new untracked files or staged changes, and it doesn’t judge functional correctness. It also checks changes left in other tracked files, not just the file just edited, so whitespace errors that existed before the work may be reported. The matcher above only reacts to Write and Edit calls; edits made via Bash commands and the like aren’t covered by this setting.

You can confirm it works in a test repository in this order:

  1. Prepare a test file tracked by Git and confirm there are no changes.
  2. Ask Claude Code to edit the file with the Edit tool and add trailing whitespace.
  3. Confirm the check message is delivered, have it remove the whitespace, and confirm the follow-up check passes.

You need to confirm both that the settings are applied and that feedback is delivered. Running the script in a terminal alone doesn’t prove it’s wired up to Claude Code.

When applying this to a project, you can extend this part: read the file path from the Hook input to check only that file, or connect the formatter or linter your project uses.

However, running the full integration test suite on every file edit makes the wait much longer. I think it’s reasonable to split the automation like this:

When Check examples
Right after a file edit Fast format checks, file-level lint
After implementing one feature Related unit tests, compilation
Before finishing the work Requirements check, related integration tests, diff review
PR CI Required checks set by the team

You can also use a Stop Hook that checks when the work ends. But if it keeps the agent working every time it fails, unnecessary loops can occur, so you need to consider state like stop_hook_active and stopping conditions. [3]

After adding checks, look at both the effect of catching errors and the extra waiting time the checks add.

6. Writing New Code Can Be Delegated the Same Way

If a bug fix is “changing wrong behavior into expected behavior,” feature development is “implementing expected behavior that doesn’t exist yet.” Both need goals and verification criteria.

6.1. Write Requirements as Verifiable Behaviors

For example, suppose we’re adding a category filter to the product list.

Add an optional categoryId filter to the product list API.

Expected behavior
- If categoryId is omitted, keep the existing query behavior.
- If categoryId is passed, return only products in that category.
- A nonexistent category returns an empty list.
- Keep the existing pagination and sorting.

How to proceed
- First read the existing query code and tests, and explain your change plan.
- Implement following the existing API response structure and project rules.
- Verify each expected behavior and update the docs.
- Ask about unclear domain rules before implementing.

Here, “a nonexistent category returns an empty list” is a requirement chosen for this example. Depending on the service, returning 404 or 400 might be right. It’s better to decide these things up front rather than leaving the agent to guess.

6.2. Extend the Verified Procedure to Feature Development

The approach verified in bug fixing can also be applied to small feature additions like this. The criteria carry over: understand the existing code first, change things based on evidence, and verify the expected behavior. On top of that, the process of interpreting new requirements and deciding the design direction is added.

Confirm that this flow works on one or two small features, and once a repeating procedure emerges, you can organize it into a feature development Skill. The procedure consists of exploring existing code, planning changes, implementation, testing, and updating docs.

Information that varies per task, like category rules or response examples, stays in the request. The common procedure goes in the Skill, and the specific requirements for this task go in the input.

For new features, passing tests isn’t always enough. For an API, you may need to check actual requests and responses; for a UI, you may need to interact with the screen. The completion report should include both what was verified and what hasn’t been checked yet.

7. Connecting Work Beyond the Code

Development work involves a lot after writing code: requesting reviews, updating docs, writing commit messages, and writing PR descriptions.

These tasks can also have clearly defined inputs and outputs.

7.1. Reflect the Final Changes in Reviews and Docs

For code review, instead of “is this code good?”, provide the final diff and requirements, and ask it to focus on possible bugs, edge cases, and changes to existing behavior. Make it include related files and the reasoning behind each comment.

PR descriptions should be written based on the actual changes rather than the initial plan. Otherwise, the direction may change during implementation while the PR still describes the old plan. Also pass along the actual execution results so that tests that weren’t run don’t end up listed as verified.

7.2. Extend with External Data and Role Splitting

If external data is needed, you can connect CLI or MCP. For example, you can pull an issue’s reproduction steps or error logs from an observability tool and use them as input for a bug fix task. Here, providing information needed for judgment — such as the time of occurrence, related requests, error messages, and reproduction conditions — rather than the entire raw log makes it easier to scope the work.

For tasks with a large exploration scope, you can also consider splitting roles into exploration, implementation, and verification. But if you use multiple agents, you need to decide who edits which files and who integrates the results. You should check that the cost of splitting roles doesn’t exceed the time it saves.

8. How Do You Know the Automation Is Working?

What I want to check is “did the agent perform the judgment and verification I consider important?” and “did that approach actually help with real work?” You need to look at both to have grounds for expanding to the next stage of automation.

8.1. Check Speed and the Basis for Judgment Together

At first, it’s good to pick a few tasks of similar size and record the following.

What to observe What it tells you
Total time to complete the task Throughput including waiting, verification, and rework
Time and reasons for human intervention Whether repeated instructions decreased, and which step gets stuck
Basis for cause diagnosis and fixes Whether changes were based on confirmed facts
Errors found again after the fix Whether verification was sufficient
Test and execution results Whether there’s evidence for calling it done
Usage cost Whether the cost is reasonable for the work saved

It’s hard to judge the effect from one successful result. You should also record tasks that stopped due to environment problems or misinterpreted requirements to see what to improve next.

8.2. Improve Procedures and Criteria Based on Verification Results

For example, if you always re-explain the DB setup before running tests, you can improve the execution docs. If it keeps getting the same exception response rule wrong, you can reinforce the rule docs and tests.

My own approach is also subject to review here. If the agent suggests a different approach that has a sound basis and solves the problem with fewer changes, I can accept it. I use my judgment criteria as a starting point, but improve them as I verify.

Conversely, if you keep adding task-specific instructions to the common rules, other tasks get more complicated too. It’s better to confirm whether a problem actually recurs before putting it in the appropriate place.

9. Designing Automation Is Also a Developer’s Job

With this setup, the amount of code a developer types directly may decrease. Instead, what you need to do becomes clearer.

Decide which problem to solve, provide the necessary context, and create criteria to check the results. When you find repeated failures, improve the work procedure or verification method.

The automation I have in mind is putting the thought process I use to look at and solve problems into the agent’s work. It’s about carrying over not just the code, but the way of understanding problems, finding evidence, making changes, and verifying results.

To do that, I need to be able to explain the judgments I usually take for granted. You can start by concretely writing down which logs you look at first, when you narrow the scope of a fix, and what you need to check before saying it’s done.

What this post treated as automation targets is code exploration, modification, running checks, and summarizing results within a defined scope. What counts as correct behavior, which design trade-offs to accept, and whether verification results are sufficient for shipping must be judged by the developer. How much you can automate depends on how clear those criteria are and how verifiable the results are.

I think automating and verifying one thing at a time like this is a good starting point. Along the way, the agent’s way of working gets refined, and the engineering I’ve been doing may become clearer too.

Thank you for reading my humble post. Insights and feedback are always welcome :)


References

  1. Claude Code — How Claude remembers your project
  2. Claude Code — Extend Claude with skills
  3. Claude Code — Hooks reference
  4. Claude Code — Automate actions with hooks
  5. Previous post — Are You Really Using AI Coding Tools Properly? (Korean)