Skip to content

kubectl cp --retries does not resume when the exec returns an error #142711

Description

@semx

What happened?

kubectl cp --retries=N resumes a copy only when the exec that runs tar ends with a nil error. TarPipe.initReadFrom runs the exec in a goroutine with cmdutil.CheckErr(t.o.execute(options)), so any error from the session exits kubectl before TarPipe.Read gets to retry; the resume only happens after the pipe was closed without an error.

Over WebSocket (the default since 1.30) a cut connection ends the exec with error reading from error stream: ... close 1006, so --retries has not resumed after a cut there. Over SPDY a cut ended the exec with nil, which is the client-go half of #142376 and the only reason --retries resumed. With the client-go change from that issue (part 1 of #142376 (comment)) a cut over SPDY is an error too, and --retries stops resuming there as well.

kubectl cp of a 3 GiB file from a pod on a stock kind v1.37.0 cluster, the connection between the apiserver and the kubelet killed 5 s in (ss -K on the node), one run per cell:

kubectl from master kubectl with the client-go change
SPDY, --retries=3 Resuming copy at 728108032 bytes, retry 1/3, complete, exit 0 error: connection closed before the command's status was received; the output may be incomplete, exit 1, no resume
WebSocket, --retries=3 error: error reading from error stream: next reader: websocket: close 1006 (abnormal closure): unexpected EOF, exit 1, no resume same

What did you expect to happen?

--retries resumes after a cut connection on either transport; that is what it was added for in #104792 (#60140).

How can we reproduce it (as minimally and precisely as possible)?

  1. A kind cluster and a pod with a large file (head -c 3221225472 /dev/zero > /tmp/big).
  2. kubectl cp --retries=3 pod:/tmp/big ./big
  3. A few seconds in, on the node: ss -K sport = :10250, which kills the apiserver's connection to the kubelet.

Anything else we need to know?

The fix is in TarPipe.initReadFrom: when MaxTries != 0, close the pipe with the exec's error instead of calling CheckErr, so that Read retries; an exit status from tar (utilexec.ExitError) stays fatal. That raises the question of which errors are worth a retry, and whether --retries=-1 needs a backoff or a stop when no byte was copied since the last attempt, since a persistent error (the pod is gone) would otherwise retry in a tight loop. I can send the PR once the client-go change has landed.

Kubernetes version

kubectl built from master (e234a6f) against kind v1.37.0 (containerd 2.3.4).

Activity

  1. added
    needs-sigIndicates an issue or PR lacks a `sig/foo` label and requires one.
    on Oct 6, 2026
  2. semx commented on Oct 6, 2026

    @semx
    MemberAuthor

    /sig cli
    /kind bug
    /area kubectl

  3. added
    needs-triageIndicates an issue or PR lacks a `triage/foo` label and requires one.
    sig/cliCategorizes an issue or PR as relevant to SIG CLI.
    kind/bugCategorizes issue or PR as related to a bug.
    and removed
    needs-sigIndicates an issue or PR lacks a `sig/foo` label and requires one.
    on Oct 6, 2026
  4. mpuckett159 commented on Oct 7, 2026

    @mpuckett159
    Contributor

    /triage accepted
    /sig apimachinery
    One caveat, we should not change the behavior with that updated error message, we should mirror what the behavior is on main for SPDY for WebSockets.

  5. kubernetes-prow commented on Oct 7, 2026

    @kubernetes-prow
    Contributor

    @mpuckett159: The label(s) sig/apimachinery cannot be applied, because the repository doesn't have them.

    Details

    In response to this:

    /triage accepted
    /sig apimachinery
    One caveat, we should not change the behavior with that updated error message, we should mirror what the behavior is on main for SPDY for WebSockets.

    Instructions for interacting with me using PR comments are available here. If you have questions or suggestions related to my behavior, please file an issue against the kubernetes-sigs/prow repository.

  6. removed
    needs-triageIndicates an issue or PR lacks a `triage/foo` label and requires one.
    on Oct 7, 2026
  7. Mujib-Ahasan commented on Oct 10, 2026

    @Mujib-Ahasan
    Contributor

    /sig api-machinery

  8. Oluwatobi-Mustapha commented on Oct 11, 2026

    @Oluwatobi-Mustapha

    @semx, thanks for the detailed report. Are you still planning to send a PR for this after the client-go change lands, or would you be comfortable with me picking it up once that prerequisite merges?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area/kubectlkind/bugCategorizes issue or PR as related to a bug.sig/api-machineryCategorizes an issue or PR as relevant to SIG API Machinery.sig/cliCategorizes an issue or PR as relevant to SIG CLI.triage/acceptedIndicates an issue or PR is ready to be actively worked on.

    Type

    No type

    Projects

    • Status
      Needs Triage

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions