Skip to main content

Step 0: Install Harbor

Follow our installation instructions to install Harbor, which involves installing the package and its dependencies.

Step 1: Create your task

Now that Harbor is installed, run the following command to create a new task directory with the required files:
This will generate a task directory with the following structure:

Step 2: Write the task instructions

Open the instruction.md file in your task directory and add the task description:
ssh-key-pair/instruction.md

Step 3: Configure task metadata

Open the task.toml file and configure your task metadata:
ssh-key-pair/task.toml
Add os = "windows" here to target Windows containers; the default is "linux". Add cpus, memory_mb, storage_mb, or gpus when the task needs explicit resources.

Step 4: Create the task environment

Open the Dockerfile in the environment/ directory that was generated:
ssh-key-pair/environment/Dockerfile
This Dockerfile defines the environment an agent will interact with through the terminal. Add any dependencies your task requires here.

Step 5: Test your solution idea

Before writing the automated solution, you’ll want to manually verify your approach works. Build and run the container interactively:
Inside the container, test that the following command solves the task without requiring interactive input:
Verify the keys were created correctly:
You should see:
Exit the container with exit or Ctrl+D.

Step 6: Write the solution script

Take the command you verified in the previous step and create the solution script. This file will be used by the Oracle agent to ensure the task is solvable. Update the solution/solve.sh file:
ssh-key-pair/solution/solve.sh
Make sure the script is executable:

Step 7: Create the test script

The test script verifies whether the agent successfully completed the task. It must produce a reward file in /logs/verifier/. This tutorial uses pytest for simplicity, but for multi-criterion verifiers, weighted scoring, or LLM-as-a-judge rubrics, consider using RewardKit instead. Update the tests/test.sh file:
ssh-key-pair/tests/test.sh
Now create the Python test file:
ssh-key-pair/tests/test_outputs.py

Step 8: Test your task with the Oracle agent

Run the following command to verify your task is solved by the solution script:
If successful, you should see output indicating the task was completed and the reward was 1.
If the Oracle agent fails, check:
  • The solution script has execute permissions
  • The Dockerfile installs all required dependencies
  • The test script correctly writes to /logs/verifier/reward.txt
  • The paths in your tests match the paths in your solution

Step 9 (Optional): Test with a real agent

Test your task with an actual AI agent to see if it can solve the task. For example, using Claude-Code with Haiku:

Step 10 (Optional): Inspect the output in the viewer

Launch the viewer to browse the run’s trajectory, verifier logs, and artifacts:
Open the latest job, pick the ssh-key-pair trial, and check the Verifier Logs tab to see your pytest output and the final reward. The Trajectory tab walks you through each step the agent took — invaluable when debugging why a real agent fails the task.

Step 11: Celebrate

Congratulations! You’ve created your first Harbor task. Your task is now ready to be used for benchmarking AI agents!