Repository navigation
kubectl cp --retries does not resume when the exec returns an error #142711
Description
Activity
- addedneeds-sigIndicates an issue or PR lacks a `sig/foo` label and requires one.Indicates an issue or PR lacks a `sig/foo` label and requires one.
on Oct 6, 2026 /sig cli
/kind bug
/area kubectl- addedneeds-triageIndicates an issue or PR lacks a `triage/foo` label and requires one.Indicates an issue or PR lacks a `triage/foo` label and requires one.sig/cliCategorizes an issue or PR as relevant to SIG CLI.Categorizes an issue or PR as relevant to SIG CLI.kind/bugCategorizes issue or PR as related to a bug.Categorizes issue or PR as related to a bug.and removedneeds-sigIndicates an issue or PR lacks a `sig/foo` label and requires one.Indicates an issue or PR lacks a `sig/foo` label and requires one.
on Oct 6, 2026 /triage accepted
/sig apimachinery
One caveat, we should not change the behavior with that updated error message, we should mirror what the behavior is on main for SPDY for WebSockets.- addedtriage/acceptedIndicates an issue or PR is ready to be actively worked on.Indicates an issue or PR is ready to be actively worked on.
on Oct 7, 2026 kubernetes-prow commented
on Oct 7, 2026 on Oct 7, 2026 – with Kubernetes ProwContributorMore actions@mpuckett159: The label(s)
sig/apimachinerycannot be applied, because the repository doesn't have them.Details
In response to this:
/triage accepted
/sig apimachinery
One caveat, we should not change the behavior with that updated error message, we should mirror what the behavior is on main for SPDY for WebSockets.Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository.
- removedneeds-triageIndicates an issue or PR lacks a `triage/foo` label and requires one.Indicates an issue or PR lacks a `triage/foo` label and requires one.
on Oct 7, 2026 /sig api-machinery
- addedsig/api-machineryCategorizes an issue or PR as relevant to SIG API Machinery.Categorizes an issue or PR as relevant to SIG API Machinery.
on Oct 10, 2026 @semx, thanks for the detailed report. Are you still planning to send a PR for this after the client-go change lands, or would you be comfortable with me picking it up once that prerequisite merges?
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsNeeds Triage
What happened?
kubectl cp --retries=Nresumes a copy only when the exec that runstarends with a nil error.TarPipe.initReadFromruns the exec in a goroutine withcmdutil.CheckErr(t.o.execute(options)), so any error from the session exits kubectl beforeTarPipe.Readgets to retry; the resume only happens after the pipe was closed without an error.Over WebSocket (the default since 1.30) a cut connection ends the exec with
error reading from error stream: ... close 1006, so--retrieshas not resumed after a cut there. Over SPDY a cut ended the exec with nil, which is the client-go half of #142376 and the only reason--retriesresumed. With the client-go change from that issue (part 1 of #142376 (comment)) a cut over SPDY is an error too, and--retriesstops resuming there as well.kubectl cpof a 3 GiB file from a pod on a stock kind v1.37.0 cluster, the connection between the apiserver and the kubelet killed 5 s in (ss -Kon the node), one run per cell:--retries=3Resuming copy at 728108032 bytes, retry 1/3, complete, exit 0error: connection closed before the command's status was received; the output may be incomplete, exit 1, no resume--retries=3error: error reading from error stream: next reader: websocket: close 1006 (abnormal closure): unexpected EOF, exit 1, no resumeWhat did you expect to happen?
--retriesresumes after a cut connection on either transport; that is what it was added for in #104792 (#60140).How can we reproduce it (as minimally and precisely as possible)?
head -c 3221225472 /dev/zero > /tmp/big).kubectl cp --retries=3 pod:/tmp/big ./bigss -K sport = :10250, which kills the apiserver's connection to the kubelet.Anything else we need to know?
The fix is in
TarPipe.initReadFrom: whenMaxTries != 0, close the pipe with the exec's error instead of callingCheckErr, so thatReadretries; an exit status fromtar(utilexec.ExitError) stays fatal. That raises the question of which errors are worth a retry, and whether--retries=-1needs a backoff or a stop when no byte was copied since the last attempt, since a persistent error (the pod is gone) would otherwise retry in a tight loop. I can send the PR once the client-go change has landed.Kubernetes version
kubectl built from master (e234a6f) against kind v1.37.0 (containerd 2.3.4).