Repository navigation
Enable exporters for resource observability, add kubelet logs in filelog receiver #509
Description
Activity
@rdimitrov I have started with first point.
For 3rd point I will add the failure modes and source details first for discussion before enabling it.@pree-dew - I think this looks good 👍 Thank you once again! 🙏
@rdimitrov Figured out the solution for all requirements here, without using any external exporter. I tried other exporters like cadvisor, kube state , metrics but was looking where with minimum moving components we can get it done.
- Using otel collector to fetch kubernetes events and ship it as log, this will help us to figure out any issue that we see with respect to deployments etc
- Using kubelet stats to get node, pod level resource utilisation metrics for cpu, memory, disk and network.
I will be raising a PR once I verify these points:
- What is the cardinality impact of including these metrics and logs
- What all failure modes can be captured without missing critical ones.
- How frequently we want to ship this.
- Any unbounded label in any of the metric
Reacted by Radoslav DimitrovPending items:
- Specify all failure modes that will be covered
- Prepare dashboard for resource metrics
- Prepare dashboard for failure modes covered by k8s events.
@domdomegg @rdimitrov Observing on staging as per the PR (#646 (comment)), so far looks good.
Cardinality is under control and k8s events have started coming as logs. I will observe it for 1 more day so that impact of deployments on cardinality is more clear with stats.
Reacted by adam jones and Radoslav Dimitrov@domdomegg @rdimitrov noticed that there is 1 deployment happened in default namespace in last 2 days, cardinality number looks under control so far.
When are we planning to do a release on production? I will make sure to keep checking till that time on regular basis.
Do you want me to send the dashboard for review or I can setup them directly on production?
@rdimitrov @domdomegg checked that production deployment gone through, everything looks okay so far. I will push small change on top of this as part of different PR :
- To change dot to underscore for ease of querying
- To reduce number of latency buckets to control cardinality.
kubelet events have started coming as logs and resource metrics are also flowing.
Reacted by adam jones and Radoslav DimitrovSounds good to me 👍
@rdimitrov @domdomegg Added all metrics and dashboard for the telemetry here : https://grafana.prod.registry.modelcontextprotocol.io/d/83a0f65c-bac4-40ce-a197-e77d67431ef4/registry-deployment-critical-issues
Let me know if anything else is required, not adding alerts as of now, want to observe the usage pattern first but if you want can add for critical events added in the dashboard.
Reacted by Radoslav DimitrovThat's awesome! 🚀 Thank you! 🙏
I think your suggestion is good and it's sensible to go with what we have now and see what alerts might be useful while we do it so we don't risk spamming ourselves 😃
@pree-dew - I think you completed everything that was supposed to be part of this issue, so feel free to close it you think so too. Also would you like to open a separate issue for analysing and adding the potential alerts?
@rdimitrov Yes, I think separate issue for alerts and these two points would be good, I will create new issue.
To change dot to underscore for ease of querying
To reduce number of latency buckets to control cardinality.Thank you! 🙏
Reacted by Radoslav DimitrovCreated new issue here #746 , closing this one.
@rdimitrov As per discussion here #478, adding todos: