YamasouA
2bed3333bc
fix lint error
2025-03-01 12:07:40 +09:00
YamasouA
75b09b4054
separete runOp
2025-03-01 11:24:23 +09:00
YamasouA
486d12efc5
call cleanup func position change
2025-03-01 00:14:49 +09:00
YamasouA
bee19638f1
tweak
2025-02-28 22:45:01 +09:00
YamasouA
038b90d475
return error instead of fatalf
2025-02-28 00:08:07 +09:00
YamasouA
f214d8e27a
delete unnecessary init
2025-02-27 00:10:24 +09:00
YamasouA
45b323d6a5
use Cleanup func
2025-02-26 23:45:42 +09:00
YamasouA
b1f6cfcfae
change defer order to pass test
2025-02-22 20:17:27 +09:00
YamasouA
fcce8aaad8
workloadExecutor's member use value not pointer
2025-02-16 23:42:20 +09:00
YamasouA
ca8a0f5f1b
separete sleep func
2025-02-13 23:46:43 +09:00
YamasouA
6d291ddc21
fix lint
2025-02-13 23:41:00 +09:00
YamasouA
a9ee6bdf81
use *e.tCtx
2025-02-13 23:33:13 +09:00
YamasouA
cc87cb54ab
delete unneccesary define
2025-02-13 23:13:57 +09:00
YamasouA
3ce36b3b3c
rename doXXX to runXXX
2025-02-13 23:11:43 +09:00
YamasouA
d202a683f5
rename workloadExecutor member name
2025-02-13 23:04:54 +09:00
YamasouA
c40e69bb4c
remove double comments
2025-02-13 23:02:41 +09:00
YamasouA
297b35873f
use workloadExecutor
2025-02-11 12:18:14 +09:00
YamasouA
479f9cd898
can pass all testcase
2025-02-11 09:26:47 +09:00
YamasouA
1b0ad78718
fix
2025-01-29 23:58:51 +09:00
YamasouA
659804b765
refactor runWorkloads
2025-01-26 19:39:38 +09:00
Kubernetes Prow Robot
bcd65ce240
Merge pull request #128667 from macsko/add_integration_tests_for_event_handling_scheduler_perf
...
Add integration tests for event handling cases in scheduler_perf
2024-12-12 13:10:26 +01:00
Kubernetes Prow Robot
ab9171b0cf
Merge pull request #129040 from sanposhiho/patch-14
...
chore: ignore dat files generated by scheduler-perf
2024-12-12 05:29:13 +00:00
Kubernetes Prow Robot
078664b424
Merge pull request #129023 from zhifei92/cleanup-actiontype
...
scheduler: Rename UpdatePodTolerations for code style consistency
2024-12-12 05:28:52 +00:00
Kensei Nakada
8f4e425daf
chore: ignore dat files generated by scheduler-perf
2024-11-30 22:23:15 +09:00
zhifei92
27608fa25d
refactor(scheduler): Rename UpdatePodTolerations for code style consistency.
2024-11-29 13:13:09 +08:00
dom4ha
67b74696f8
Adjust performance test threshold limits
2024-11-25 15:07:15 +00:00
Patrick Ohly
ac3d43a8a6
scheduler_perf: work around incorrect gotestsum failure reports
...
Because Go does not a "pass" action for
benchmarks (https://github.com/golang/go/issues/66825#issuecomment-2343229005 ),
gotestsum reports a successful benchmark run as failed
(https://github.com/gotestyourself/gotestsum/issues/413#issuecomment-2343206787 ).
We can work around that in each benchmark and sub-benchmark by emitting the
output line that `go test` expects on stdout from the test binary for success.
2024-11-18 12:35:05 +01:00
Patrick Ohly
369a18a3a1
scheduler_perf: simplify flags, fix output
...
The "disabled by label filter" message for benchmarks printed the pointer to
the filter string, not the filter string itself. This mistake gets avoided and
the code becomes simpler when not using pointers.
2024-11-18 12:32:59 +01:00
Maciej Skoczeń
de8e8c5404
Add integration tests for event handling cases in scheduler_perf
2024-11-13 13:17:48 +00:00
Kubernetes Prow Robot
8115baca00
Merge pull request #128666 from macsko/fix_scale_down_in_eventhandlingpodupdate_scheduler_perf_test_case
...
Fix pod scale down failure in EventHandlingPodUpdate scheduler_perf test
2024-11-12 16:28:47 +00:00
Kubernetes Prow Robot
fb033826a8
Merge pull request #128170 from sanposhiho/async-preemption
...
feature(KEP-4832): asynchronous preemption
2024-11-07 19:44:54 +00:00
Maciej Skoczeń
379bff8dc9
Fix pod scale down failure in EventHandlingPodUpdate scheduler_perf test case
2024-11-07 13:48:50 +00:00
Patrick Ohly
0301b6b504
scheduler_perf: fix steady-state pod creation/deletion
...
This fixes an issue in
TestSchedulerPerf/SteadyStateClusterResourceClaimTemplate:
scheduler_perf.go:1542: FATAL ERROR: op 7: delete scheduled pods: client rate limiter Wait returned an error: rate: Wait(n=1) would exceed context deadline
That occurs when the test is almost done, but hasn't observed all scheduled
pods yet. The previous attempt to address this error wasn't actually 100%
correct. It covered the case when the context has already been canceled, but
not this particular "will reach deadline soon".
2024-11-07 09:36:36 +01:00
Kensei Nakada
4a084d54d2
feat: set the threashold on the scheduler-perf test case
2024-11-07 14:09:35 +09:00
Kensei Nakada
4b92f6d398
fix the broken part due to the merge
2024-11-07 14:09:35 +09:00
Kensei Nakada
69a8d0ec0b
feature(KEP-4832): asynchronous preemption
2024-11-07 14:09:34 +09:00
Patrick Ohly
30f5282656
DRA API: rename DeviceCapacity.Quantity to DeviceCapacity.Value
...
Based on review
feedback (https://github.com/kubernetes/kubernetes/pull/127511#discussion_r1823521172 ).
2024-11-06 13:03:20 +01:00
Patrick Ohly
33ea278c51
DRA: use v1beta1 API
...
No code is left which depends on the v1alpha3, except of course the code
implementing that version.
2024-11-06 13:03:19 +01:00
Kubernetes Prow Robot
0fad78930f
Merge pull request #127904 from towca/jtuznik/dra-autoscaling
...
DRA: allow Cluster Autoscaler to integrate with DRA scheduler plugin
2024-11-06 10:01:29 +00:00
Kubernetes Prow Robot
9bbb46d05f
Merge pull request #128566 from macsko/run_scheduler_perf_with_queueinghints_enabled_disabled
...
Run scheduler_perf with QueueingHints both enabled and disabled
2024-11-05 14:53:29 +00:00
Kuba Tużnik
8d489425aa
scheduler/dynamicresources: extract obtaining and tracking in-memory modifications of DRA objects
...
All logic related to obtaining DRA objects and tracking modifications
to ResourceClaims in-memory is extracted to DefaultDRAManager, which
implements framework.SharedDRAManager.
This is intended to be a no-op in terms of the DRA plugin behavior.
2024-11-05 14:11:04 +01:00
Kubernetes Prow Robot
2bb886ce2a
Merge pull request #128482 from sanposhiho/scheduler-perf-ff
...
fix: register QHint metrics only when available
2024-11-05 12:15:30 +00:00
Kubernetes Prow Robot
c69f150008
Merge pull request #127277 from pohly/dra-structured-performance
...
kube-scheduler: enhance performance for DRA structured parameters
2024-11-05 10:05:29 +00:00
Kensei Nakada
0bf95100f1
fix: register QHint metrics only when available
2024-11-05 18:52:27 +09:00
Maciej Skoczeń
e44041ee47
Run scheduler_perf with QueueingHints both enabled and disabled
2024-11-05 09:13:03 +00:00
Patrick Ohly
7863d9a381
DRA scheduler: refactor CEL compilation cache
...
A better place is the cel package because a) the name can become shorter
and b) it is tightly coupled with the compiler there.
Moving the compilation into the cache simplifies the callers.
2024-11-05 08:34:42 +01:00
Maciej Skoczeń
8371a35824
Split scheduler_perf config into subdirectories
2024-11-04 08:45:34 +00:00
Patrick Ohly
bc55e82621
DRA scheduler: maintain a set of allocated device IDs
...
Reacting to events from the informer cache (indirectly, through the assume
cache) is more efficient than repeatedly listing it's content and then
converting to IDs with unique strings.
goos: linux
goarch: amd64
pkg: k8s.io/kubernetes/test/integration/scheduler_perf
cpu: Intel(R) Core(TM) i9-7980XE CPU @ 2.60GHz
│ before │ after │
│ SchedulingThroughput/Average │ SchedulingThroughput/Average vs base │
PerfScheduling/SchedulingWithResourceClaimTemplateStructured/5000pods_500nodes-36 54.70 ± 6% 76.81 ± 6% +40.42% (p=0.002 n=6)
PerfScheduling/SteadyStateClusterResourceClaimTemplateStructured/empty_100nodes-36 106.4 ± 4% 105.6 ± 2% ~ (p=0.413 n=6)
PerfScheduling/SteadyStateClusterResourceClaimTemplateStructured/empty_500nodes-36 120.0 ± 4% 118.9 ± 7% ~ (p=0.117 n=6)
PerfScheduling/SteadyStateClusterResourceClaimTemplateStructured/half_100nodes-36 112.5 ± 4% 105.9 ± 4% -5.87% (p=0.002 n=6)
PerfScheduling/SteadyStateClusterResourceClaimTemplateStructured/half_500nodes-36 87.13 ± 4% 123.55 ± 4% +41.80% (p=0.002 n=6)
PerfScheduling/SteadyStateClusterResourceClaimTemplateStructured/full_100nodes-36 113.4 ± 2% 103.3 ± 2% -8.95% (p=0.002 n=6)
PerfScheduling/SteadyStateClusterResourceClaimTemplateStructured/full_500nodes-36 65.55 ± 3% 121.30 ± 3% +85.05% (p=0.002 n=6)
geomean 90.81 106.8 +17.57%
2024-11-01 13:23:06 +01:00
Patrick Ohly
814c9428fd
DRA scheduler: cache compiled CEL expressions
...
DeviceClasses and different requests are very likely to contain the same
expression string. We don't need to compile that over and over again.
To avoid hanging onto that cache longer than necessary, it's currently tied to
each PreFilter/Filter combination. It might make sense to move this up into the
scheduler plugin and thus reuse compiled expressions for different pods.
goos: linux
goarch: amd64
pkg: k8s.io/kubernetes/test/integration/scheduler_perf
cpu: Intel(R) Core(TM) i9-7980XE CPU @ 2.60GHz
│ before │ after │
│ SchedulingThroughput/Average │ SchedulingThroughput/Average vs base │
PerfScheduling/SchedulingWithResourceClaimTemplateStructured/5000pods_500nodes-36 33.95 ± 4% 36.65 ± 2% +7.95% (p=0.002 n=6)
PerfScheduling/SteadyStateClusterResourceClaimTemplateStructured/empty_100nodes-36 105.8 ± 2% 106.7 ± 3% ~ (p=0.177 n=6)
PerfScheduling/SteadyStateClusterResourceClaimTemplateStructured/empty_500nodes-36 100.7 ± 1% 119.7 ± 3% +18.82% (p=0.002 n=6)
PerfScheduling/SteadyStateClusterResourceClaimTemplateStructured/half_100nodes-36 90.78 ± 1% 121.10 ± 4% +33.40% (p=0.002 n=6)
PerfScheduling/SteadyStateClusterResourceClaimTemplateStructured/half_500nodes-36 50.51 ± 7% 63.72 ± 3% +26.17% (p=0.002 n=6)
PerfScheduling/SteadyStateClusterResourceClaimTemplateStructured/full_100nodes-36 103.7 ± 5% 110.2 ± 2% +6.32% (p=0.002 n=6)
PerfScheduling/SteadyStateClusterResourceClaimTemplateStructured/full_500nodes-36 28.50 ± 2% 28.16 ± 5% ~ (p=0.102 n=6)
geomean 64.99 73.15 +12.56%
2024-11-01 13:20:06 +01:00
Patrick Ohly
941d17b3b8
DRA scheduler: code cleanups
...
Looking up the slice can be avoided by storing it when allocating a device.
The AllocationResult struct is small enough that it can be copied by value.
goos: linux
goarch: amd64
pkg: k8s.io/kubernetes/test/integration/scheduler_perf
cpu: Intel(R) Core(TM) i9-7980XE CPU @ 2.60GHz
│ before │ after │
│ SchedulingThroughput/Average │ SchedulingThroughput/Average vs base │
PerfScheduling/SchedulingWithResourceClaimTemplateStructured/5000pods_500nodes-36 33.30 ± 2% 33.95 ± 4% ~ (p=0.288 n=6)
PerfScheduling/SteadyStateClusterResourceClaimTemplateStructured/empty_100nodes-36 105.3 ± 2% 105.8 ± 2% ~ (p=0.524 n=6)
PerfScheduling/SteadyStateClusterResourceClaimTemplateStructured/empty_500nodes-36 100.8 ± 1% 100.7 ± 1% ~ (p=0.738 n=6)
PerfScheduling/SteadyStateClusterResourceClaimTemplateStructured/half_100nodes-36 90.96 ± 2% 90.78 ± 1% ~ (p=0.952 n=6)
PerfScheduling/SteadyStateClusterResourceClaimTemplateStructured/half_500nodes-36 49.84 ± 4% 50.51 ± 7% ~ (p=0.485 n=6)
PerfScheduling/SteadyStateClusterResourceClaimTemplateStructured/full_100nodes-36 103.8 ± 1% 103.7 ± 5% ~ (p=0.582 n=6)
PerfScheduling/SteadyStateClusterResourceClaimTemplateStructured/full_500nodes-36 27.21 ± 7% 28.50 ± 2% ~ (p=0.065 n=6)
geomean 64.26 64.99 +1.14%
2024-11-01 13:19:51 +01:00