Skip to content

Latest commit

 

History

2 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 

Repository files navigation

Production-Grade Incremental Kubernetes Adoption for 40+ Java Microservices

Core Achievement Successfully migrated an existing distributed Java microservices system (40+ services) to Kubernetes in a phased, zero-downtime manner — without disrupting ongoing operations.

Key Innovation:

  • Designed a hybrid co-existence architecture where containerized (new/updated) services and non-containerized (legacy/running) services operate as one unified system
  • Enabled service-by-service migration — gradually containerizing one microservice at a time
  • Implemented fine-grained traffic control with percentage-based routing (e.g., 5% → 30% → 100% to new container version)
  • Maintained single codebase, single branch, same configs for both container and non-container deployments
  • Developer experience: unchanged — just build Docker image → Helm one-click deploy to any env (dev/test/prod)

Started with pure Kubernetes bridging (Ingress + custom routing), later enhanced with Istio for advanced traffic shifting, resilience, and observability.

This approach is far beyond standard "lift-and-shift" — it required deep expertise in architecture, development, and DevOps to bridge environments seamlessly at scale.

Notes:

  • This diagram is a high-level illustration intended to demonstrate bi-directional interactions between in-cluster and out-of-cluster services.
  • In production, the system includes extensive customization based on business requirements (such as custom gateways, synchronization mechanisms, and control logic). These details are part of internal company implementations and are not disclosed in this repository.
  • This approach allows us to migrate services incrementally, control traffic by percentage, and shift traffic to containerized versions service by service, achieving zero downtime and avoiding large, one-shot cutovers.

img.png

Production Snippets & Examples (Anonymized)

Selected excerpts from internal docs and simplified configs (details abstracted for confidentiality).

1. Helm Values Snippet (Deployment Example)

#  Helm values excerpt
{{- if .Values.ingress }}
{{- if .Values.ingress.enabled -}}
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
{{- with .Values.ingressAnnotations }}
annotations:
{{- toYaml . | nindent 4 }}
{{- end }}
{{- $serviceName := .Values.serviceName }}
name: {{ $serviceName }}-ingress
namespace: {{ .Values.namespace }}
spec:
rules:
- http:
  paths:
  {{- range .Values.ingressPaths }}
  {{- with . }}
    - backend:
      service:
      name: {{ $serviceName }}
      port:
      number: 80
      path: {{ .path }}
      pathType: {{ .type | default "Prefix" }}
      {{- end }}
      {{- end }}
      {{- end }}
      {{- end }}

2. Deployment YAML Excerpt (Real Production Example)

apiVersion: apps/v1
kind: Deployment
metadata:
  annotations:
    deployment.kubernetes.io/revision: '22'
    meta.helm.sh/release-name: company
    meta.helm.sh/release-namespace: prod
  generation: 25
  labels:
    app: app-1
    app.kubernetes.io/managed-by: Helm
    k8s.eip.work/layer: svc
    version: v1
  name: app-1
  namespace: prod
  resourceVersion: '68392681'
  uid: eb7b0406-46cd-47d5-9dc7-37dad9ed2afa
spec:
  progressDeadlineSeconds: 600
  replicas: 2
  revisionHistoryLimit: 10
  selector:
    matchLabels:
      app: app-1
      version: v1
  strategy:
    rollingUpdate:
      maxSurge: 50%
      maxUnavailable: 25%
    type: RollingUpdate
  template:
    metadata:
      annotations:
        tag: HMicFdrY1VIBDUq25dPk11
      creationTimestamp: null
      labels:
        app: app-1
        version: v1
    spec:
      affinity:
        podAntiAffinity:
          requiredDuringSchedulingIgnoredDuringExecution:
            - labelSelector:
                matchExpressions:
                  - key: app
                    operator: In
                    values:
                      - app-1
              topologyKey: kubernetes.io/hostname
      containers:
        - args:
            - '-jar'
            - '-Dserver.port=8080'
            - '-Drun.platform=k8s-cloud'
            - '-javaagent:/usr/local/agent/skywalking-agent.jar'
            - '-Dskywalking.collector.backend_service=x.x.x.x:11801'
            - '-Dskywalking.agent.service_name=app-1'
            - '-Xms3072m'
            - '-Xmx3072m'
            - '-XX:+UseG1GC'
            - '-XX:+UnlockExperimentalVMOptions'
            - '-XX:G1ReservePercent=15'
            - '-XX:G1HeapRegionSize=4'
            - '-XX:G1HeapWastePercent=10'
            - '-XX:G1MixedGCLiveThresholdPercent=80'
            - '-XX:G1NewSizePercent=45'
            - '-XX:G1MaxNewSizePercent=55'
            - '-XX:InitiatingHeapOccupancyPercent=50'
            - '-XX:-UseBiasedLocking'
            - /app-1.jar
            - '--apollo_meta=http://x.x.x.x:8088'
            - '--apollo_cluster=default'
          command:
            - java
          envFrom:
            - configMapRef:
                name: javaapp-base
          image: 'xx.com/prod/app-1:1.0.0'
          imagePullPolicy: Always
          lifecycle:
            preStop:
              exec:
                command:
                  - curl
                  - '-XPOST'
                  - '127.0.0.1:9090/management/shutdown'
          livenessProbe:
            failureThreshold: 3
            initialDelaySeconds: 90
            periodSeconds: 30
            successThreshold: 1
            tcpSocket:
              port: 8080
            timeoutSeconds: 5
          name: app-1
          ports:
            - containerPort: 8080
              name: http
              protocol: TCP
            - containerPort: 9090
              name: management
              protocol: TCP
          readinessProbe:
            failureThreshold: 3
            httpGet:
              path: /management/health
              port: 9090
              scheme: HTTP
            initialDelaySeconds: 90
            periodSeconds: 30
            successThreshold: 1
            timeoutSeconds: 5
          resources:
            limits:
              cpu: '2'
              memory: 4840Mi
            requests:
              cpu: 200m
              memory: 2816Mi
          terminationMessagePath: /dev/termination-log
          terminationMessagePolicy: File
          volumeMounts:
            - mountPath: /logs
              name: logs
      dnsPolicy: ClusterFirst
      restartPolicy: Always
      schedulerName: default-scheduler
      securityContext: {}
      terminationGracePeriodSeconds: 30
      tolerations:
        - effect: NoExecute
          key: node.kubernetes.io/not-ready
          operator: Exists
          tolerationSeconds: 10
        - effect: NoExecute
          key: node.kubernetes.io/unreachable
          operator: Exists
          tolerationSeconds: 10
      volumes:
        - emptyDir:
            sizeLimit: 10Gi
          name: logs

About

real-world cloud-native configurations and practices

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors