Skip to samples
Typed Decision Bench · tasks · Real-time and agents

k8s-issue-kind-routing

agent routing and skill selection · devops · 200 items · primary question type choice. Source open-index/open-github-issues (ODC-BY-1.0); labels are outcome labels, never an LLM judge. Download these samples (JSON).

SystemDecisionScoreAccuracy (top pick)Answered
Jev77.572.5%200 / 200
decider-2b75.466.0%200 / 200
OpenJev73.871.0%200 / 200
JevFish73.662.0%200 / 200
System One Scorer73.561.0%200 / 200
NanoJev61.534.0%200 / 200
Sample 1 of 6 · real-time-and-agents:k8s-115483 · 181 state tokens (common tier)

Turn 1 — the request (identical for every system)

State (JSON)
{
  "repo": "kubernetes/kubernetes",
  "title": "Secrets does not creating",
  "body": "### What happened? I am writing to report an issue with the latest version of Kubernetes in Azure, version v1.24.6. Before this version, when creating a Kubernetes cluster, the service account and secrets were automatically created and everything worked smoothly. However, with version v1.24.6, I have encountered an error when attempting to connect the cluster to Azure DevOps. Also during execution command in version v1.23.8 \"kubectl create serviceaccount some_name\", security token generating automatically. But not in v1.24.6. I found only 2 secrets in new cluster v1.24.6 comparing with v1.23.8 where was more than 30 secrets. I am checking all namespaces with command \"kubectl get secrets …"
}

Question kind · choice · primary (ranked)

Read `repo`, `title` and `body`. Which kubernetes issue kind will maintainers assign to this issue?

7 options
  • bug — Something is broken or behaves contrary to documentation (kind/bug)
  • feature — Request for new functionality or an enhancement (kind/feature)
  • cleanup — Refactor, tech-debt removal, deprecation of old code, no user-facing change (kind/cleanup)
  • documentation — Docs are missing, wrong or unclear (kind/documentation)
  • support — A usage question or help request rather than a defect (kind/support)
  • failing_test — A test fails consistently in CI (kind/failing-test)
  • flake — A test fails intermittently / flaky CI job (kind/flake)

Turn 2 — each system's response · Turn 3 — the grade

Expected answer: support

SystemTop pickGradeP(expected)Proper scoreDistribution
Jevbugwrong23.0%0.407
  • bug 77.0%
  • support 23.0%
  • cleanup 0.0%
  • documentation 0.0%
  • failing_test 0.0%
  • flake 0.0%
decider-2bbugwrong7.9%0.258
  • bug 79.5%
  • support 7.9%
  • feature 4.2%
  • documentation 3.7%
  • failing_test 2.4%
  • cleanup 1.7%
OpenJevbugwrong0.0%0.016
  • bug 98.4%
  • feature 1.6%
  • flake 0.0%
  • cleanup 0.0%
  • support 0.0%
  • documentation 0.0%
JevFishbugwrong9.9%0.259
  • bug 81.8%
  • support 9.9%
  • feature 2.7%
  • cleanup 2.7%
  • documentation 2.1%
  • flake 0.7%
System One Scorersupportcorrect38.6%0.737
  • support 38.6%
  • bug 36.9%
  • cleanup 8.4%
  • documentation 4.9%
  • failing_test 4.3%
  • flake 3.7%
NanoJevfailing_testwrong8.2%0.481
  • failing_test 28.9%
  • bug 23.3%
  • feature 18.4%
  • flake 14.2%
  • support 8.2%
  • cleanup 4.9%
Sample 2 of 6 · real-time-and-agents:k8s-115556 · 253 state tokens (common tier)

Turn 1 — the request (identical for every system)

State (JSON)
{
  "repo": "kubernetes/kubernetes",
  "title": "[Flaky Test] Conformance - GCE - master - kubetest2",
  "body": "### Which jobs were failing? Conformance - GCE - master - kubetest2 ### Which tests failed? It looks like the overall build may be failing ### When has it been failing? From 1:35pm - 3:35pm ### Testgrid link https://k8s-testgrid.appspot.com/sig-release-master-blocking#Conformance%20-%20GCE%20-%20master%20-%20kubetest2 ### Reason for failure (if possible) It looks like build/release.sh script blew up: ``` DOCKER_CLI_EXPERIMENTAL=enabled docker buildx build --load -t kube-build:build-6454b08013-5-v1.27.0-go1.20-bullseye.0 --pull=false --build-arg=KUBE_CROSS_IMAGE=registry.k8s.io/build-image/kube-cross --build-arg=KUBE_CROSS_VERSION=v1.27.0-go1.20-bullseye.0 …"
}

Question kind · choice · primary (ranked)

Read `repo`, `title` and `body`. Which kubernetes issue kind will maintainers assign to this issue?

7 options
  • bug — Something is broken or behaves contrary to documentation (kind/bug)
  • feature — Request for new functionality or an enhancement (kind/feature)
  • cleanup — Refactor, tech-debt removal, deprecation of old code, no user-facing change (kind/cleanup)
  • documentation — Docs are missing, wrong or unclear (kind/documentation)
  • support — A usage question or help request rather than a defect (kind/support)
  • failing_test — A test fails consistently in CI (kind/failing-test)
  • flake — A test fails intermittently / flaky CI job (kind/flake)

Turn 2 — each system's response · Turn 3 — the grade

Expected answer: failing_test

SystemTop pickGradeP(expected)Proper scoreDistribution
Jevflakewrong1.0%0.020
  • flake 99.0%
  • failing_test 1.0%
  • cleanup 0.0%
  • support 0.0%
  • documentation 0.0%
  • feature 0.0%
decider-2bflakewrong10.6%0.214
  • flake 87.9%
  • failing_test 10.6%
  • bug 1.1%
  • cleanup 0.1%
  • feature 0.1%
  • documentation 0.1%
OpenJevflakewrong0.2%0.007
  • flake 99.6%
  • failing_test 0.2%
  • feature 0.1%
  • bug 0.1%
  • documentation 0.0%
  • cleanup 0.0%
JevFishflakewrong26.7%0.542
  • flake 61.0%
  • failing_test 26.7%
  • cleanup 7.1%
  • bug 3.8%
  • support 0.7%
  • feature 0.5%
System One Scorerflakewrong9.3%0.409
  • flake 57.7%
  • cleanup 10.5%
  • bug 9.3%
  • failing_test 9.3%
  • support 6.9%
  • documentation 3.5%
NanoJevflakewrong25.6%0.656
  • flake 28.5%
  • failing_test 25.6%
  • bug 17.0%
  • feature 11.7%
  • support 7.0%
  • cleanup 5.4%
Sample 3 of 6 · real-time-and-agents:k8s-122585 · 221 state tokens (common tier)

Turn 1 — the request (identical for every system)

State (JSON)
{
  "repo": "kubernetes/kubernetes",
  "title": "kubelet creates two duplicate containers",
  "body": "### What happened? In our k8s cluster, I created a pod, but kubelet pulled up two containers,as shown below: ![pod](https://github.com/kubernetes/kubernetes/assets/62041640/be6c78cf-9241-4273-8896-588c84f04664) ![container](https://github.com/kubernetes/kubernetes/assets/62041640/628f4ad5-ca48-4cbf-afdb-de55eb40a63e) Judging from the kubelet log, a timeout occurred when kubelet created the first container.kubelet may think container creation failed, then kubelet created a second container. When the second container is created successfully, the first container is also created successfully, so container duplication occurs. ### What did you expect to happen? I think kubelet should not create …"
}

Question kind · choice · primary (ranked)

Read `repo`, `title` and `body`. Which kubernetes issue kind will maintainers assign to this issue?

7 options
  • bug — Something is broken or behaves contrary to documentation (kind/bug)
  • feature — Request for new functionality or an enhancement (kind/feature)
  • cleanup — Refactor, tech-debt removal, deprecation of old code, no user-facing change (kind/cleanup)
  • documentation — Docs are missing, wrong or unclear (kind/documentation)
  • support — A usage question or help request rather than a defect (kind/support)
  • failing_test — A test fails consistently in CI (kind/failing-test)
  • flake — A test fails intermittently / flaky CI job (kind/flake)

Turn 2 — each system's response · Turn 3 — the grade

Expected answer: bug

SystemTop pickGradeP(expected)Proper scoreDistribution
Jevbugcorrect100.0%1.000
  • bug 100.0%
  • documentation 0.0%
  • feature 0.0%
  • cleanup 0.0%
  • support 0.0%
  • failing_test 0.0%
decider-2bbugcorrect81.9%0.981
  • bug 81.9%
  • support 4.2%
  • failing_test 4.0%
  • feature 3.3%
  • documentation 3.0%
  • cleanup 2.5%
OpenJevbugcorrect99.1%1.000
  • bug 99.1%
  • feature 0.9%
  • documentation 0.0%
  • flake 0.0%
  • support 0.0%
  • cleanup 0.0%
JevFishbugcorrect85.1%0.986
  • bug 85.1%
  • cleanup 4.4%
  • support 4.0%
  • feature 3.0%
  • failing_test 1.5%
  • documentation 1.0%
System One Scorerbugcorrect66.3%0.930
  • bug 66.3%
  • support 10.8%
  • cleanup 9.2%
  • flake 5.6%
  • documentation 3.2%
  • feature 2.6%
NanoJevfeaturewrong16.2%0.583
  • feature 21.1%
  • failing_test 19.7%
  • bug 16.2%
  • flake 13.1%
  • support 12.8%
  • cleanup 10.6%
Sample 4 of 6 · real-time-and-agents:k8s-126119 · 176 state tokens (common tier)

Turn 1 — the request (identical for every system)

State (JSON)
{
  "repo": "kubernetes/kubernetes",
  "title": "Do not start cadvisor when feature PodAndContainerStatsFromCRI is enabled",
  "body": "### What would you like to be added? Remove the need to start cadvisor, which has a periodic housekeeping task. That leaves on-demand invocations to cadvisor APIs ### Why is this needed? Cadvisor housekeeping task is a visible overhead that collects host OS stats, which can be entirely replaced by CRI APIs. The only use left for having housekeeping task is to [collect rootfs stats](https://github.com/kubernetes/kubernetes/blob/master/pkg/kubelet/cadvisor/cadvisor_linux.go#L151), which can be changed to an on-demand call. Also see prior issues that attempted to manage the overhead of cadvisor: https://github.com/kubernetes/kubernetes/pull/124520 …"
}

Question kind · choice · primary (ranked)

Read `repo`, `title` and `body`. Which kubernetes issue kind will maintainers assign to this issue?

7 options
  • bug — Something is broken or behaves contrary to documentation (kind/bug)
  • feature — Request for new functionality or an enhancement (kind/feature)
  • cleanup — Refactor, tech-debt removal, deprecation of old code, no user-facing change (kind/cleanup)
  • documentation — Docs are missing, wrong or unclear (kind/documentation)
  • support — A usage question or help request rather than a defect (kind/support)
  • failing_test — A test fails consistently in CI (kind/failing-test)
  • flake — A test fails intermittently / flaky CI job (kind/flake)

Turn 2 — each system's response · Turn 3 — the grade

Expected answer: feature

SystemTop pickGradeP(expected)Proper scoreDistribution
Jevfeaturecorrect87.0%0.984
  • feature 87.0%
  • cleanup 12.0%
  • bug 1.0%
  • support 0.0%
  • documentation 0.0%
  • failing_test 0.0%
decider-2bcleanupwrong10.9%0.266
  • cleanup 82.1%
  • feature 10.9%
  • documentation 2.6%
  • bug 1.9%
  • failing_test 1.3%
  • support 0.8%
OpenJevfeaturecorrect99.8%1.000
  • feature 99.8%
  • bug 0.1%
  • flake 0.1%
  • support 0.0%
  • documentation 0.0%
  • cleanup 0.0%
JevFishcleanupwrong45.1%0.746
  • cleanup 45.2%
  • feature 45.1%
  • support 3.8%
  • bug 3.6%
  • documentation 1.8%
  • flake 0.5%
System One Scorercleanupwrong19.7%0.505
  • cleanup 57.7%
  • feature 19.7%
  • bug 7.6%
  • support 5.9%
  • documentation 3.5%
  • flake 3.5%
NanoJevfeaturecorrect33.2%0.723
  • feature 33.2%
  • failing_test 25.1%
  • flake 13.4%
  • bug 13.0%
  • support 8.3%
  • cleanup 5.1%
Sample 5 of 6 · real-time-and-agents:k8s-129759 · 183 state tokens (common tier)

Turn 1 — the request (identical for every system)

State (JSON)
{
  "repo": "kubernetes/kubernetes",
  "title": "Documentation of pod selection for node-pressure eviction is confusing regarding QoS",
  "body": "In [the Node-pressure doc explaining the ranking of pods for eviction](https://kubernetes.io/docs/concepts/scheduling-eviction/node-pressure-eviction/#pod-selection-for-kubelet-eviction), there is a note saying that QoS are not used for this ranking (which appears to be correct, AFAICT from the code). However, this section still mentions QoS extensively, and it reads a bit ambiguous. In particular, this section: >As a result, kubelet ranks and evicts pods in the following order: > > 1. BestEffort or Burstable pods where the usage exceeds requests. These pods are evicted based on their Priority and then by how much their usage level exceeds the request. > 2. Guaranteed pods and Burstable …"
}

Question kind · choice · primary (ranked)

Read `repo`, `title` and `body`. Which kubernetes issue kind will maintainers assign to this issue?

7 options
  • bug — Something is broken or behaves contrary to documentation (kind/bug)
  • feature — Request for new functionality or an enhancement (kind/feature)
  • cleanup — Refactor, tech-debt removal, deprecation of old code, no user-facing change (kind/cleanup)
  • documentation — Docs are missing, wrong or unclear (kind/documentation)
  • support — A usage question or help request rather than a defect (kind/support)
  • failing_test — A test fails consistently in CI (kind/failing-test)
  • flake — A test fails intermittently / flaky CI job (kind/flake)

Turn 2 — each system's response · Turn 3 — the grade

Expected answer: documentation

SystemTop pickGradeP(expected)Proper scoreDistribution
Jevdocumentationcorrect100.0%1.000
  • documentation 100.0%
  • flake 0.0%
  • cleanup 0.0%
  • support 0.0%
  • failing_test 0.0%
  • feature 0.0%
decider-2bdocumentationcorrect87.5%0.988
  • documentation 87.5%
  • bug 9.6%
  • cleanup 1.5%
  • feature 0.5%
  • support 0.3%
  • failing_test 0.3%
OpenJevdocumentationcorrect99.6%1.000
  • documentation 99.6%
  • bug 0.1%
  • support 0.1%
  • feature 0.1%
  • flake 0.1%
  • cleanup 0.0%
JevFishdocumentationcorrect78.6%0.961
  • documentation 78.6%
  • bug 18.1%
  • cleanup 1.6%
  • support 1.1%
  • failing_test 0.3%
  • feature 0.3%
System One Scorerdocumentationcorrect59.3%0.895
  • documentation 59.3%
  • bug 12.8%
  • support 12.2%
  • cleanup 11.1%
  • feature 2.0%
  • flake 1.7%
NanoJevdocumentationcorrect49.3%0.847
  • documentation 49.3%
  • failing_test 12.3%
  • support 12.0%
  • feature 8.7%
  • bug 8.0%
  • flake 6.8%
Sample 6 of 6 · real-time-and-agents:k8s-129802 · 288 state tokens (common tier)

Turn 1 — the request (identical for every system)

State (JSON)
{
  "repo": "kubernetes/kubernetes",
  "title": "[Flaking Test] [sig-api-machinery] k8s.io/kubernetes/test/integration/apiserver/coordinatedleaderelection.coordinatedleaderelection",
  "body": "### Which jobs are flaking? master-blocking - integration-master ### Which tests are flaking? k8s.io/kubernetes/test/integration/apiserver/coordinatedleaderelection.coordinatedleaderelection [Prow](https://prow.k8s.io/view/gs/kubernetes-ci-logs/logs/ci-kubernetes-integration-master/1881787128904421376) [Triage](https://storage.googleapis.com/k8s-triage/index.html?test=k8s.io%2Fkubernetes%2Ftest%2Fintegration%2Fapiserver%2Fcoordinatedleaderelection.coordinatedleaderelection&xjob=e2e-kops) ### Since when has it been flaking? [1/14/2025, 12:41:26 AM](https://prow.k8s.io/view/gs/kubernetes-ci-logs/logs/ci-kubernetes-integration-master/1878882224963588096) [1/18/2025, 1:29:23 …"
}

Question kind · choice · primary (ranked)

Read `repo`, `title` and `body`. Which kubernetes issue kind will maintainers assign to this issue?

7 options
  • bug — Something is broken or behaves contrary to documentation (kind/bug)
  • feature — Request for new functionality or an enhancement (kind/feature)
  • cleanup — Refactor, tech-debt removal, deprecation of old code, no user-facing change (kind/cleanup)
  • documentation — Docs are missing, wrong or unclear (kind/documentation)
  • support — A usage question or help request rather than a defect (kind/support)
  • failing_test — A test fails consistently in CI (kind/failing-test)
  • flake — A test fails intermittently / flaky CI job (kind/flake)

Turn 2 — each system's response · Turn 3 — the grade

Expected answer: flake

SystemTop pickGradeP(expected)Proper scoreDistribution
Jevflakecorrect100.0%1.000
  • flake 100.0%
  • feature 0.0%
  • cleanup 0.0%
  • support 0.0%
  • documentation 0.0%
  • failing_test 0.0%
decider-2bflakecorrect87.9%0.988
  • flake 87.9%
  • failing_test 9.6%
  • bug 1.8%
  • cleanup 0.3%
  • feature 0.2%
  • documentation 0.1%
OpenJevflakecorrect99.9%1.000
  • flake 99.9%
  • failing_test 0.1%
  • bug 0.0%
  • feature 0.0%
  • documentation 0.0%
  • cleanup 0.0%
JevFishflakecorrect67.8%0.920
  • flake 67.8%
  • failing_test 23.0%
  • cleanup 4.8%
  • bug 2.1%
  • feature 1.2%
  • support 0.9%
System One Scorerflakecorrect77.6%0.969
  • flake 77.6%
  • failing_test 7.6%
  • bug 5.4%
  • cleanup 4.4%
  • support 2.2%
  • documentation 1.4%
NanoJevflakecorrect25.9%0.669
  • flake 25.9%
  • failing_test 20.9%
  • feature 16.5%
  • bug 15.3%
  • support 10.3%
  • cleanup 7.7%