mirror of
https://hubproxy.babadafafafafa.cn/https://github.com/kube-vip/kube-vip.git
synced 2026-09-20 08:03:47 +08:00
* fix(e2e/services): pass metrics address as container arg The prometheusHTTPServer env var is never read by kube-vip; environment parsing only knows prometheus_server and ignores empty values, so every parallel DaemonSet still tried to bind the :2112 flag default on the shared host network. Pass the address as --prometheusHTTPServer instead, where an empty value really disables the metrics server. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Maximilian Rink <maximilian.rink@telekom.de> * fix(egress): restore CIDR fallback and cluster-wide SNAT exclusion Auto-discovery dropped the kube-controller-manager fallback, breaking clusters whose CNI does its own IPAM without Node PodCIDR allocation. Fall back to the --cluster-cidr/--service-cluster-ip-range flags when the ServiceCIDR or Node APIs yield nothing. Per-node PodCIDR discovery combined with getSameFamilyCidr also narrowed the SNAT exclusion to the local node's pod CIDR, so egress traffic to pods on other nodes was SNAT'd to the VIP. Exclude every same-family pod CIDR in both the nftables and iptables paths. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Maximilian Rink <maximilian.rink@telekom.de> * test(e2e/services): fail failover tests when the VIP never recovers leaderFailover and podFailover returned nil when the watch drained without any post-kill update passing httpTest, so an unreachable VIP after failover passed CI. Keep the warn-and-continue polling, but turn the drain path into an error. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Maximilian Rink <maximilian.rink@telekom.de> * test(e2e/services): require sustained quiet window in lease error check A single quiet 2 s interval can coincide with the pause between failing election retries and produce a false pass for the #1664 guard. Require three consecutive intervals without counter growth and reset the streak whenever the counter grows. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Maximilian Rink <maximilian.rink@telekom.de> * fix(svcs): log service deletion only when a tracked service is removed deleteTrackedService logged '(svcs) deleted' for every watch event of every non-LoadBalancer service, even when nothing was tracked. Move the log inside the svcCtx guard. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Maximilian Rink <maximilian.rink@telekom.de> * fix(svcs): skip API refresh when egress endpoints are unchanged updateEgressConfiguration issued a Services Get on every no-op watch event before short-circuiting. Compare the ActiveEndpoint annotations first and only refresh the service from the API when they differ. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Maximilian Rink <maximilian.rink@telekom.de> * test(e2e/services): drop dead StartServiceTest and restore run summary StartServiceTest was dead code once main.go inlined the phases, and the E2E_KEEP_LOGS-conditional temp-dir cleanup plus the 'Testing Complete' summary silently disappeared with it. Reinstate both in run() and collect every test result so parallel failures are all reported instead of collapsing to the errgroup's first error. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Maximilian Rink <maximilian.rink@telekom.de> * style(e2e): gofmt test files gofmt -l flagged e2e_bgp_healthcheck_test.go (the Mode == ModeBGP block body was never re-indented) and a trailing space in e2e_sigusr1_test.go. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Maximilian Rink <maximilian.rink@telekom.de> * ci: write ginkgo JSON reports where the artifact upload finds them Without --output-dir ginkgo resolves the report path relative to the suite directory, so /tmp/kube-vip-test-report-*.json landed under testing/e2e/tmp/ and the /tmp/kube-vip-test* artifact glob missed it. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Maximilian Rink <maximilian.rink@telekom.de> * ci: scope cancel-in-progress to PRs and restore tag builds cancel-in-progress: true also cancelled in-progress main builds on back-to-back merges; limit cancellation to pull_request runs. The branches: [main] filter added in the refactor dropped the tag pushes the bare push trigger previously covered, and the release workflow only publishes images, so restore lint/test coverage for v* tags. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Maximilian Rink <maximilian.rink@telekom.de> * ci: fail the validation job on unformatted Go code make check only formats '*.go pkg cmd' and merely warns, which let unformatted files under testing/ land. Add a blocking gofmt -l check over the whole repo that prints the offending files. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Maximilian Rink <maximilian.rink@telekom.de> * test(e2e/services): kill kube-vip pods in the test namespace during failover Since the parallel-phase refactor the kube-vip DaemonSet runs in the per-test namespace, but the failover killer still listed pods in kube-system, so it never deleted a single pod and no failover was ever induced. The lenient drain path masked this; the strict assertion exposed it. Target the test namespace, gate success on events observed after at least one actual pod kill, extend the watch window to 60s to leave convergence margin after the ~25s kill churn, and return distinct errors for aborted watch, zero kills, and unreachable VIP. Detach test cleanup from the errgroup context (context.WithoutCancel) and demote cleanup-path Fatals to errors so one failing test no longer kills sibling teardown. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Maximilian Rink <maximilian.rink@telekom.de> * test(e2e/bgp): raise route-destination timeout for parallel runs With ginkgo --procs=4, sibling processes create kind clusters and load images on the same runner; the 30s window for the first route check overlaps that contention and flakes (observed >13s convergence even in passing specs). Use 120s, matching the BGP health-check suite's re-announcement timeout. Eventually returns on success, so passing specs are unaffected. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Maximilian Rink <maximilian.rink@telekom.de> --------- Signed-off-by: Maximilian Rink <maximilian.rink@telekom.de> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>