Kubernetes Reference
Commands, customization, environment variables, and troubleshooting for the Kubernetes template.
The scripts install the Helm release agentos into the agentos namespace. Override with AGENTOS_RELEASE and AGENTOS_NAMESPACE.
Manage
| Task | Command |
|---|---|
| Roll to a new image tag | IMAGE_TAG=v2 ./scripts/k8s/redeploy.sh (build and push the tag first) |
| Restart pods in place | ./scripts/k8s/redeploy.sh |
| Sync supported env variables | ./scripts/k8s/env-sync.sh (updates nonempty values from a fixed allowlist that includes RUNTIME_ENV; defaults to .env.production, or pass .env) |
| Tail logs | kubectl logs deploy/agentos -n agentos -f |
| Port-forward the API | kubectl port-forward svc/agentos 8000:8000 -n agentos |
| Roll back a release | helm rollback agentos -n agentos |
| Tear down | ./scripts/k8s/down.sh (add --yes to skip the confirmation) |
Production auth
Token-Based Authorization is on by default. Production startup requires JWT_VERIFICATION_KEY or a readable JWKS file at the pod path in JWT_JWKS_FILE; otherwise the process exits.
Token-Based Auth gives you three things:
- Protected runtime access. Protected API routes require a valid credential. Health, discovery, and API documentation remain public; Slack and MCP OAuth have their own authentication flows.
- Per-request identity. Middleware validates the token and exposes its
user_id, optionalsession_id, scopes, and claims to the request. - Scope-based permissions. Token scopes control access to AgentOS routes and resources.
The template already sets AuthorizationConfig(user_isolation=True). Authenticated non-admin REST access is scoped to the principal; local dev mode disables scope enforcement and is open when no credentials are configured. This does not scope Platform Manager’s direct database tools to that REST identity. See User Isolation for the boundaries and admin exceptions.
To disable JWT authentication in a private deployment with another auth layer, set authorization=False, remove JWT_VERIFICATION_KEY and JWT_JWKS_FILE from the running service, and rebuild or redeploy the app as needed. Disabling scope enforcement alone does not remove JWT validation while those credentials remain configured. MCP OAuth remains enabled while MCP_CONNECT_SECRET is set.
Customize
Ask your coding agent to run /create-agent, or do it by hand. Create agents/my_agent.py:
from agno.agent import Agent
from app.learning import shared_learning
from app.settings import default_model
from db import get_postgres_db
INSTRUCTIONS = """\
What the agent does, which tools it uses, the rules to follow when answering.
"""
my_agent = Agent(
id="my-agent",
name="My Agent",
user_id="anonymous-user", # Local fallback; authenticated runs supply identity.
model=default_model(),
db=get_postgres_db(),
instructions=INSTRUCTIONS,
learning=shared_learning,
add_datetime_to_context=True,
add_history_to_context=True,
num_history_runs=5,
)The fallback user ID lets local anonymous calls use the shared learning machine; those calls share one profile. Authenticated run identity overrides this default.
Import it in app/main.py and update the agents argument in the existing AgentOS call. Keep its other arguments, including teams, workflows, knowledge, and registry:
from agents.my_agent import my_agent
agent_os = AgentOS(
# Keep the other arguments from the existing call.
agents=[platform_builder, platform_manager, platform_engineer, my_agent],
)Add its UI metadata beneath the existing manifest: key in app/config.yaml:
my-agent:
description: "What the agent does."
quick_prompts:
- "First example prompt"
- "Second example prompt"
- "Third example prompt"Local containers reload Python source changes. After editing app/config.yaml, run docker compose restart agentos-api to reload the manifest. For production, build and push a new image tag, then run IMAGE_TAG=<tag> ./scripts/k8s/redeploy.sh. If the release still runs the official image, point it at your registry first: IMAGE_REPOSITORY=<registry>/agentos IMAGE_TAG=<tag> ./scripts/k8s/up.sh.
app/settings.py defines default_model(), used by every agent. Change it in one place:
from agno.models.anthropic import Claude
def default_model():
return Claude(id="claude-sonnet-5")Add anthropic to pyproject.toml, set the provider key in your env, and regenerate pins:
./scripts/generate_requirements.shRebuild locally with docker compose up -d --build. For production, build and push a new tag, then roll to it:
docker build -t <registry>/agentos:v2 . && docker push <registry>/agentos:v2
IMAGE_TAG=v2 ./scripts/k8s/redeploy.shenv-sync.sh uses a fixed allowlist that includes RUNTIME_ENV, AGENTOS_URL, the JWT_JWKS_FILE path, the template's supported secrets, and DB_PASS. It does not deliver the referenced JWKS file or sync a new provider key such as ANTHROPIC_API_KEY. Provide the JWKS file through a custom image or chart volume. Deliver a new provider key via extraEnv and helm upgrade.
Agno ships 100+ toolkits. See Toolkits.
from agno.tools.slack import SlackTools
my_agent = Agent(
# Keep the agent’s existing configuration.
tools=[SlackTools()],
)- Edit
pyproject.toml. - Regenerate pins:
./scripts/generate_requirements.sh(addupgradeto refresh every pin). - Rebuild locally with
docker compose up -d --build, or build and push a new tag and roll to it withIMAGE_TAG=<tag> ./scripts/k8s/redeploy.sh.
Set both variables in your env file:
SLACK_BOT_TOKEN=xoxb-...
SLACK_SIGNING_SECRET=...Sync with ./scripts/k8s/env-sync.sh. The interface activates automatically and routes messages to the Agno team. Change team= in app/main.py to select another team, or replace it with agent=my_agent to route to an agent. See Slack setup.
The deployment check runs daily by default (ENABLE_DEPLOY_CHECK=True); it uses fixed checks without model calls. The run-evals schedule is always registered but starts disabled because it uses model calls. Enable it from the AgentOS UI. Both workflows remain runnable on demand. Startup reapplies ENABLE_DEPLOY_CHECK to the deployment-check schedule; the enabled state of an existing run-evals schedule is preserved.
In the cluster these are chart values. Set ENABLE_DEPLOY_CHECK and EVALS_* via extraEnv and helm upgrade; env-sync.sh does not sync them. Enable the registered run-evals schedule from the AgentOS UI.
Format, validate, and run evals
Run evals against a dedicated local test platform with no concurrent writers. The starter’s cleanup hooks remove components and learning state created during a case; concurrent application writes can be removed too. The same prerequisite applies to scheduled evals. See eval setup and isolation.
The host scripts require uv. The setup script creates a Python 3.14 venv:
./scripts/venv_setup.sh
source .venv/bin/activate| Task | Command |
|---|---|
| Format | ./scripts/format.sh |
| Lint and type-check | ./scripts/validate.sh |
| Run smoke evals | python -m evals --tag smoke |
./scripts/mcp_check.sh runs inside the container, so it needs no venv.
Environment variables
| Variable | Required | Default | Description |
|---|---|---|---|
OPENAI_API_KEY | Yes | - | Models and embeddings. |
RUNTIME_ENV | No | prd | dev disables scope enforcement; configured JWT credentials still enable token validation. Compose sets it for local. Keep production on prd so scope enforcement remains enabled. |
JWT_VERIFICATION_KEY | Production | - | Public key from os.agno.com. Quote the value so the multi-line PEM parses as one variable. |
JWT_JWKS_FILE | Production | - | Path inside the pod to a JWKS file. up.sh and env-sync.sh set only the jwtJwksFile path. The current chart does not mount the file. |
MCP_CONNECT_SECRET | No | - | OAuth consent secret (16+ chars) for connecting claude.ai and ChatGPT to /mcp. up.sh generates it into .env.production when the deploy has a public URL (INGRESS_HOST or AGENTOS_URL); set it by hand otherwise. |
AGENTOS_MCP_SIGNING_KEY | No | generated | Optional high-entropy signing-key material (32+ chars) for OAuth tokens. Unset, a strong key is generated and persisted in the database. Rotating it invalidates outstanding tokens. |
AGENTOS_URL | No | http://agentos:8000 | Scheduler base URL. The chart resolves an explicit value first, then the ingress URL, then the release service URL. Set it only for a custom domain or tunnel. When MCP_CONNECT_SECRET is set, OAuth metadata uses this URL as its public origin. |
ENABLE_DEPLOY_CHECK | No | True | Daily deployment-check cron. |
EVALS_TAG | No | smoke | Eval tag the run-evals workflow runs. |
EVALS_CASE_TIMEOUT_SECONDS | No | 90 | Fallback timeout for cases without an explicit timeout. |
EVALS_SUITE_TIMEOUT_SECONDS | No | derived from selected cases | Sum of selected case timeouts plus 30 seconds per case, with a 60-second floor. A positive integer overrides this ceiling. |
PARALLEL_API_KEY | No | - | WebSearch uses the Parallel SDK when set, keyless MCP otherwise. |
SLACK_BOT_TOKEN | No | - | Set with the signing secret to enable Slack. |
SLACK_SIGNING_SECRET | No | - | Set with the bot token to enable Slack. |
DB_HOST / DB_PORT / DB_USER / DB_PASS / DB_DATABASE | No | matches compose | Postgres connection. up.sh generates DB_PASS once and saves it to your env file. |
DB_DRIVER | No | postgresql+psycopg | SQLAlchemy driver. |
AGNO_DEBUG | No | False | Verbose Agno logs. Compose sets it for dev. |
WAIT_FOR_DB | No | True in Helm | If True, the entrypoint blocks on the database before starting. The Helm chart and Compose set it to True. |
AGENTOS_NAMESPACE | No | agentos | Namespace the k8s scripts target. |
AGENTOS_RELEASE | No | agentos | Helm release name the k8s scripts target. |
IMAGE_REPOSITORY | No | agnohq/agentos | Image the chart deploys. Read by up.sh. |
IMAGE_TAG | No | latest | Image tag. up.sh installs it; redeploy.sh rolls the release to it. |
IMAGE_PULL_POLICY | No | IfNotPresent | Set Never for images loaded into kind. Read by up.sh. |
INGRESS_HOST | No | - | Publishes the API behind your ingress controller at this host. Read by up.sh. |
INGRESS_CLASS | No | - | Ingress class name, for example nginx. Read by up.sh. |
Troubleshooting
up.sh deploys into your current kubectl context and verifies it can reach the cluster first. Point kubectl at the target cluster and confirm kubectl get namespace works, then rerun.
Expected. At os.agno.com, choose Connect OS → Live, enter your AgentOS URL, name it Live AgentOS, turn on Token-Based Authorization (JWT) on the connection panel, and connect. The UI generates the public key. If the OS is already connected, enable the setting under Settings → OS & Security. Paste the full PEM into the script prompt. To add a PEM later, set JWT_VERIFICATION_KEY and run ./scripts/k8s/env-sync.sh. To use JWKS, first provide the file through a custom image or chart volume, then set JWT_JWKS_FILE to its pod path and sync.
Non-dev mode requires JWT verification configuration. Configured JWT credentials also enable validation in dev mode. Set JWT_VERIFICATION_KEY and sync. JWT_JWKS_FILE works only when a custom image or chart mount already provides a readable file at that pod path. For a private deployment using another auth layer, follow the credential-removal steps under Production auth.
The cluster can't pull the image. Confirm the tag was pushed and the cluster has access to your registry; for private registries, set imagePullSecrets in charts/agentos/values.yaml. On kind, kind load docker-image the tag and deploy with IMAGE_PULL_POLICY=Never.
The Postgres volume reads its password only on first initialization, so a lost or regenerated DB_PASS locks the app out of an existing volume. Restore the DB_PASS that up.sh saved to your env file and sync, fix the database in place with ALTER USER, or delete the PVC to reinitialize. Deleting the PVC deletes all data.
AGENTOS_URL resolves automatically: explicit value, then ingress URL, then in-cluster service DNS. If you set it by hand, make sure the pod can reach that URL, then run ./scripts/k8s/env-sync.sh.
up.sh generates MCP_CONNECT_SECRET only when the deploy has a public URL (INGRESS_HOST or an explicit AGENTOS_URL). Deployed without one? Set MCP_CONNECT_SECRET and a public AGENTOS_URL in .env.production and run ./scripts/k8s/env-sync.sh.