Recently, while browsing Google Cloud's developer environment setup documentation, I noticed something new amidst the usual gcloud CLI, Cloud Shell, and Cloud Workstations:
agy plugin install https://github.com/google/skills/plugins/cloud/google-cloud-developer
agy is the Google Antigravity CLI. But what really made me stop wasn't this new CLI, but the URL following it: Google has started packaging its own product knowledge as agent plugins and publishing them on GitHub.
As it turns out, when I checked my own Claude Code, this plugin was already installed on my machine. So, in this post, we won't bother installing the Antigravity CLI first; instead, we'll use what's already available to see what it actually does.
What's inside this plugin?
Let's take it apart. The entire package is actually quite small: 5 skills, 1 set of routing rules, and 1 MCP server, totaling 1549 lines of markdown with no code at all:
skills/gcloud/ # Safety guardrails and syntax validation for gcloud commands (272 lines + 340 lines of reference docs)
skills/retrieving-developer-knowledge/ # Querying official documentation (103 lines + 182 lines of reference docs)
skills/finding-google-skills/ # Fetching skills from a remote directory on demand (134 lines)
skills/google-cloud-recipe-onboarding/ # Onboarding for a user's first project (228 lines)
skills/google-cloud-recipe-auth/ # Credentials and ADC selection (260 lines)
rules/google-cloud-discovery.md # Skill routing table (30 lines)
Interestingly, four sets of manifests are packed into the same directory:
.claude-plugin/plugin.json # Claude Code
.codex-plugin/plugin.json # Codex
gemini-extension.json # Gemini CLI / Antigravity
plugin.json # Universal format
The universal plugin.json starts like this:
{
"$schema": "https://agent-plugins.org/schemas/1.0.0/plugin.schema.json",
"name": "google-cloud-developer",
"version": "1.1.2",
"author": { "name": "Google LLC" },
"license": "Apache-2.0"
}
agent-plugins.org is a cross-tool plugin schema. Google didn't create a proprietary format for its own Antigravity; instead, they made one package that four different types of agents can consume. In other words, you can use this content even without installing the Antigravity CLI.
Installation: Three Paths to Choose From
Antigravity CLI (As per documentation)
agy plugin install https://github.com/google/skills/plugins/cloud/google-cloud-developer
Claude Code
claude plugin marketplace add google/skills
claude plugin install google-cloud-developer@google-plugins
Universal
npx skills add google/skills
After installation, you need to enable the API:
gcloud services enable developerknowledge.googleapis.com --project=YOUR_PROJECT_ID
To confirm where it's installed:
$ ls ~/.claude/plugins/cache/google-plugins/google-cloud-developer/
1.1.2
Test 1: Can guardrails stop hallucinated flags?
The core of this plugin is the gcloud skill. Its opening paragraph is very blunt:
All pre-existing knowledge of
gcloudcommands, flags, flag values, and positional argument syntax is stale and prone to hallucination.
In plain English: The gcloud syntax you (the model) remember is outdated; do not write it from memory. Consequently, it sets a hard rule: before taking action, you must run gcloud help <leaf command>. Furthermore, parent-level help doesn't count; you must validate down to the lowest-level sub-command. It even explicitly forbids using web searches to check syntax; gcloud help is the sole authority.
Is this rule necessary? Let's test it with Cloud Run's scaling flags:
$ for f in --min-scale --max-scale --min-instances --max-instances; do
gcloud help run deploy | rg -q -- "$f" && echo "$f exists" || echo "$f does not exist"
done
--min-scale does not exist
--max-scale does not exist
--min-instances exists
--max-instances exists
--min-scale / --max-scale were the syntax for Knative and early Cloud Run, but they are no longer used. However, these strings are everywhere in old articles and Stack Overflow answers online, so models are very likely to spit out a command that looks reasonable but fails immediately upon execution. By running gcloud help first, this problem disappears.
By the way, gcloud run deploy doesn't actually have a --dry-run flag; it only has --async. The skill's rule is: "If the help output lists --dry-run or --validate-only, you must run it once first." If it's listed, run it; if not, skip it.
Test 2: How much context does a blind list consume?
There is a rule in the skill that I initially thought was a bit tedious:
DO NOT execute any
listcommand without including at least one data reduction flag (--limit,--filter, or--format).
A quick measurement shows why. I have a project running 52 Cloud Run services:
# A. Blind list
$ gcloud run services list --project=YOUR_PROJECT_ID --format=json | wc -c
281918
# B. Projecting only the two required fields (still 52 services, none missing)
$ gcloud run services list --project=YOUR_PROJECT_ID \
--format="json(metadata.name,status.url)" | wc -c
7963
281,918 bytes vs. 7,963 bytes—a 35x difference. In terms of tokens, that's roughly 70k vs. 2k. In other words, a list command without --format would consume 70k tokens, and 97% of what was cut out consists of annotations, conditions, revision templates, and a bunch of timestamps.
The skill also teaches a very practical "probe schema before querying full data" trick:
gcloud <GROUP> <RESOURCE> list --limit=1 --format=json
First, grab one entry to see what the JSON looks like. After confirming the field paths, then construct the --filter and --format.
Other execution rules follow the same logic:
- Run only one command at a time; no
&&chaining. - Prohibit pipes,
$( ), and redirection (the reason being to keep commands readable for human review). - Always add
--project=<PROJECT_ID>; do not rely on the default values of the active config. - Always add
--quietto avoid getting stuck on interactive confirmations in environments without a TTY.
Test 3: Which commands is it forbidden to run?
This is the section I find most worth copying. The skill contains an explicit denylist of operations that must never be executed automatically without explicit human authorization:
- Any changes to IAM policy / role / binding (risk of privilege escalation or locking oneself out).
-
gcloud * delete(irreversible resource destruction). -
gcloud billing *(risk of billing spikes). -
gcloud organizations *(organization-level settings affect everyone). -
gcloud kms *(potential to permanently lock data). -
gcloud infra-manager deployments apply(automatically running IaC might delete a whole batch of resources). - Proactively enabling APIs is also on the list, the reason being that enabling an API might trigger billing.
I find that last one particularly interesting. Enabling an API sounds harmless, but the skill explicitly states "assume necessary APIs are enabled"; if they need to be enabled, go back and ask the human. This level of conservatism shows they have seriously thought about what happens when an agent is connected to a production environment.
Test 4: Developer Knowledge—Querying docs instead of memory
The plugin includes an MCP server:
{
"mcpServers": {
"developer-knowledge": {
"type": "streamable-http",
"url": "https://developerknowledge.googleapis.com/mcp"
}
}
}
Here I encountered the first pitfall. Just because an MCP server is declared in the configuration file doesn't mean it actually connects. The tools listed on my end were only authenticate and complete_authentication; the answer_query and search_documents mentioned in the documentation didn't appear at all, meaning it was stuck in the authentication process.
Interestingly, the skill itself had already anticipated this, stating in black and white:
A declared server is not always a connected server. Some clients cannot complete the MCP handshake with this server and expose no
answer_query,search_documentsorget_documentstool at all, even though the plugin declares one.
It then provides a REST fallback. You can use the gcloud credentials you already have to make the call, without needing to apply for a separate API key:
curl -s -X POST "https://developerknowledge.googleapis.com/v1:answerQuery" \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "X-Goog-User-Project: $(gcloud config get-value project)" \
-H "Content-Type: application/json" \
-d '{"query": "How do I deploy a container to Cloud Run with gcloud?"}'
HTTP 200, and the returned data is more sophisticated than I expected:
$ jq '.answer | keys' response.json
["answerText", "citations", "references"]
$ jq '.answer.references | length' response.json
10
$ jq -r '.answer.references[].documentReference.documentChunk.document.uri' response.json | sort -u
https://docs.cloud.google.com/run/docs/deploying
https://docs.cloud.google.com/run/docs/quickstarts/deploy-container
https://docs.cloud.google.com/sdk/gcloud/reference/run/deploy
https://docs.cloud.google.com/run/docs/configuring/services/containers
https://docs.cloud.google.com/run/docs/create-jobs
...
It doesn't just provide an answer; it attaches 10 official documents as sources, and citations use character ranges to map to the reference index:
{"startIndex": 447, "endIndex": 505, "sources": [{"referenceIndex": 3}, {"referenceIndex": 7}]}
Characters 447 to 505 in the answer come from the 3rd and 7th documents. With this level of granularity in citations, verifying the truth of the answer becomes much easier.
To query precise syntax, use another endpoint. I tested it with Cloud Run IAM permission strings:
curl -s -G "https://developerknowledge.googleapis.com/v1/documents:searchDocumentChunks" \
--data-urlencode "query=cloud run invoker IAM permission run.routes.invoke" \
--data-urlencode "pageSize=3" \
-H "Authorization: Bearer $(gcloud auth print-access-token)" \
-H "X-Goog-User-Project: $(gcloud config get-value project)"
The retrieved permission strings:
run.routes.invoke
run.services.create
run.services.delete
run.services.get
run.services.list
run.services.update
These service.resource.verb format strings are where models are most likely to miss a character or add an extra 's'. Checking them is much more reliable than guessing.
A small pitfall to note: the skill's own fallback documentation says the response should be read from documentChunks[].content, but the actual field returned is results[].content. This minor discrepancy is telling in itself. Even official documentation can be out of sync with the actual API, which is why the "verify before acting" rule is indeed useful.
Test 5: How 137 skills avoid consuming context
finding-google-skills solves another problem: Google's skill directory is massive, but it's impossible to preload everything.
Its approach is to keep only a single page routing table and fetch the remote directory only when actually needed:
$ curl -sSL https://raw.githubusercontent.com/google/skills/main/index.json -o index.json
$ jq '.skills | length' index.json
137
$ wc -c index.json
84476
137 skills, 84KB. Stuffing it all into the context is too wasteful, so the skill teaches you to filter with jq before reading:
$ curl -sSL https://raw.githubusercontent.com/google/skills/main/index.json \
| jq -r '.skills[] | select((.name+" "+.description)|test("cloud run";"i")) | .name'
cloud-run-basics
cloud-run-alert-configuration
cloud-monitoring-metric-selection
google-cloud-global-frontend-configuration
...
Select up to three of the most relevant ones, and only then fetch their entrypoints. This "remote directory, load on demand" pattern is, in my opinion, the most reasonable solution for handling a large number of skills currently.
One more detail: this skill explicitly forbids using curl -k to bypass certificate verification when authentication fails, even specifically naming PowerShell's ServicePointManager callback. The reason given is excellent: "You are about to execute commands based on what you fetch; an unverified directory is worse than no directory at all."
Five Practical Scenarios
Scenario 1: Connecting an agent to production, but only allowing it to look, not touch
This is the one I wanted most. Previously, I didn't quite dare point Claude Code at a production project because it might suddenly decide to "clean up" a batch of resources. With the denylist, the boundaries become very clear.
Help me check which Cloud Run services in the production project haven't had a new revision in over three months.
It will list, filter, and organize the data into a table for you, but it will stop and ask when it encounters a delete operation. What it saves is the time spent flipping through documents and assembling commands; the decision-making power remains in your hands. This boundary is very well-defined.
Scenario 2: Starting a GCP project from scratch and deploying a LINE Bot to Cloud Run
I want to start a new GCP project to deploy a LINE Bot. Help me go through the process from creating the project, linking billing, enabling APIs, to deploying to Cloud Run.
google-cloud-recipe-onboarding will guide you through the entire sequence: billing → project → enabling APIs → deployment. It verifies syntax before each command and will definitely stop to ask you before enabling APIs (because it's on the denylist). This is perfect for people who only start a new project once in a long while and have to re-google everything each time.
Scenario 3: Credential Hell—Works locally, but 403 on Cloud Run
This is probably the most common "hitting a wall" scenario in GCP.
My service runs fine locally, but after deploying to Cloud Run, calling Vertex AI keeps returning a 403. Help me clarify which credential is the issue.
google-cloud-recipe-auth is responsible for clarifying the relationships and applicable scenarios for user credentials, ADC, and service accounts, while Developer Knowledge is responsible for looking up the precise IAM permission strings. Combined, these two are much faster than digging through five-year-old answers on Stack Overflow.
Scenario 4: Cloud Resource Inventory and Billing Health Check
List the names, URLs, and last deployment times of all Cloud Run services in this project, and identify those that might no longer be in use.
The key here is that 35x difference measured earlier. Without guardrails, the raw JSON of 52 services would be enough to fill the agent's context with junk data before it even begins the actual analysis. With --format projection and the habit of "probing schema with --limit=1 first," 52 services only consume 2k tokens, allowing it to handle several times more resources.
Scenario 5: Avoiding the model's old memories when asking about new features
For products that iterate very quickly like Gemini API, Cloud Run, and Firebase, the model's training data is bound to be outdated.
How do I set min instances and startup CPU boost in Cloud Run now?
With Developer Knowledge, the answer will include links to official documents and character-level citation ranges. More importantly, if it can't find the information, it will say so directly:
Presenting recalled documentation as a retrieved result is the worst available outcome, because nothing in the reply distinguishes it from a real lookup.
"I didn't find it; this is an answer from memory" is far more valuable than a flag that sounds smooth but is incorrect.
Summary
From my testing, the things that left the deepest impression on me were actually two things unrelated to functionality.
First, the entire package is 1549 lines of markdown, without a single line of code. All security is implemented through text-based rules: verify syntax first, limit output, explicit denylist, and honest reporting of failure. This means you can also use the same method to write your own team's runbooks and security rules as skills: the barrier to entry is much lower than imagined.
Second, four sets of manifests are squeezed into the same directory. Google didn't create a proprietary format for Antigravity; instead, they made the same package usable across Claude Code, Codex, and Gemini CLI. Skills are becoming a common format that isn't tied to a specific tool.
If you have a GCP project and are using a coding agent, this plugin takes about five minutes to set up. Just for the fact that it won't hallucinate flags or touch IAM on its own, it's well worth it.

Top comments (0)