Logo

Personal Ops Runbook

Personal runbook covering infrastructure operations for Cloud, Kubernetes, OpenStack, and Ceph environments. Includes deployment and teardown procedures, node management, cluster monitoring setup, and incident response workflows compiled from day-to-day operational work. Intended strictly for personal reference — configurations and scripts are environment-specific and not guaranteed to work as-is elsewhere.

Dynamic Rebalancing

Nguồn: Lightbits Private Cloud Administration Guide (3.19.x / 3.20.x)

Fail in Place

Khi chế độ fail in place được kích hoạt, Lightbits cluster sẽ cố gắng di chuyển các replica của volume từ các node bị lỗi sang các node khỏe mạnh khác, đồng thời vẫn tuân thủ các yêu cầu về failure domain.

Chế độ fail in place được kích hoạt theo yêu cầu của người dùng. Khoảng thời gian từ lúc node gặp sự cố cho đến khi cluster bắt đầu phục hồi volume được xác định bởi tham số cấu hình cluster tên là DurationToTurnIntoPermanentFailure. Sau khi hết thời gian được định nghĩa trong biến DurationToTurnIntoPermanentFailure, quá trình nhân bản mới bắt đầu hoạt động.

Ví dụ về Fail in Place

Kích hoạt chế độ fail in place bằng cách bật feature-flag:

lbcli enable feature-flag fail-in-place [flags]

Ví dụ:

enable a feature flag in place with given feature flag name

lbcli --jwt $JWT enable feature-flag fail-in-place

Tắt chế độ fail in place bằng cách vô hiệu hóa feature-flag:

lbcli disable feature-flag fail-in-place [flags]

Ví dụ:

disable a feature flag in place with given feature flag name

lbcli --jwt $JWT disable feature-flag fail-in-place

Ví dụ về cách đặt thời gian chờ trước khi cluster rebalance:

lbcli -J $JWT update cluster-config --parameter=DurationToTurnIntoPermanentFailure --value=20m