Russian documentation: README.ru.md, SMOKE.ru.md
HMS Proxy is a catalog-aware Hive Metastore federation and compatibility proxy for mixed
Apache Hive 3.1.3, Hive 4.1.x, and Hortonworks Data Platform 3.1.0.x environments.
It gives you one (or several) production-facing HMS Thrift endpoint(s) that can federate catalogs across multiple backend metastores, bridge Apache 3.1.3 / HDP 3.1.0.x / Hive 4.1.x API differences in either direction, and establish a clear security boundary between clients and backend HMS services.
Use one production-facing HMS Thrift endpoint for HiveServer2 and direct HMS API clients (or
several, on different ports, when clients of different Hive versions cannot share a single
front door — Thrift has no version-negotiation handshake), while routing requests to multiple
backend metastores by explicit catName or by legacy catalog<separator>db database names.
This lets you centralize catalog-aware routing and selective exposure without forcing clients to understand backend metastore layout.
Expose an Apache Hive Metastore 3.1.3, Hortonworks 3.1.0.x or Hive 4.1.x front door (one
primary listener plus optional additional listeners on separate ports) and connect it to any
mix of Apache 3.1.3, Hortonworks 3.1.0.x, or Hive 4.1.x backends. The proxy downgrades
selected *_req APIs for older Hortonworks backends, upgrades the two positional read methods
Hive 4 removed (get_table, get_table_objects_by_name) when fronting a Hive 4 backend with
an Apache 3.1.3 client, and downgrades Hive 4-only read/standard-DDL *_req wrappers to their
positional Apache 3.1.3 equivalents when serving Hive 4 clients against an Apache 3.1.3
backend. Hortonworks-specific RPCs are exposed via the HDP standalone-metastore jar when a
matching HDP backend is configured; truly Hive 4-only APIs (data connectors, scheduled
queries, stored procedures, packages, ACID v2 extensions) respond with
TApplicationException UNKNOWN_METHOD.
This makes the proxy a practical bridge for mixed Apache/HDP/Hive 4 estates and staged migrations in either direction, not just a request router.
Run the proxy as the security boundary between clients and backend metastores with optional Kerberos/SASL on the front door, optional outbound Kerberos to backends, and optional impersonation of the authenticated Kerberos caller.
This keeps authentication, identity propagation, and backend exposure policy concentrated in one place.
Two naming layers appear throughout this README:
| Term | Meaning |
|---|---|
| External name | The client-facing proxy namespace. Examples: catName=catalog2, catalogName=catalog2, legacy dbName=catalog2__sales. |
| Internal name | The real names on the selected backend HMS. Example: backend catName=hive, dbName=sales. |
| Default catalog | routing.default-catalog, used when a request does not carry a trustworthy catalog hint. |
| Ambiguous request | The request carries conflicting catalog hints, or a multi-catalog write does not resolve to exactly one catalog. |
Routing is then decided like this:
| Request shape | How the proxy picks a backend | What happens without namespace |
|---|---|---|
| Object-scoped reads and writes | Prefer explicit proxy catName. Otherwise parse external dbName / fullTableName such as catalog2__sales. Otherwise use routing.default-catalog for compatibility. |
Routed to routing.default-catalog. |
| Session-level and global reads | Use routing.default-catalog. |
Routed to routing.default-catalog. |
| ACID RPCs whose payload still names an object | Route by that payload, for example dbName or fullTableName in get_valid_write_ids, allocate_table_write_ids, compact, or add_dynamic_partitions. |
If no routable object namespace remains, fall through to the next row. |
| Txn / lock lifecycle RPCs that only carry ids | Pin to routing.default-catalog, for example open_txns, commit_txn, abort_txn, check_lock, unlock, and heartbeat. |
Routed to routing.default-catalog; non-ACID SELECT, eligible NO_TXN DDL and non-transactional write locks on non-default catalogs can use the synthetic shim. |
| Global writes and catalog-registry changes | Require exactly one owned namespace in a multi-catalog deployment. | If that ownership is ambiguous, the proxy fails safely instead of guessing a target catalog. |
If a request carries a backend catName such as hive instead of a proxy catalog id, the proxy
treats that field as non-authoritative for compatibility and falls back to dbName or
routing.default-catalog.
For mutations, the proxy prefers policy-driven guarantees over best-effort guessing: deterministic routing, explicit namespace ownership, no silent split-brain writes, and safe failure when a mutation stays ambiguous.
These switches change client-visible names or SQL text, not backend selection:
| Switch | Effect |
|---|---|
routing.catalog-db-separator |
Changes the external legacy spelling, for example catalog2__sales instead of catalog2.sales. |
federation.preserve-backend-catalog-name=true |
Returns backend catName / catalogName such as hive, but routing still follows the external dbName or explicit proxy catalog. |
federation.view-text-rewrite.mode=REWRITE |
Rewrites view SQL between external and internal names; it does not change backend selection for the RPC itself. |
| RPC group | Status | Behavior |
|---|---|---|
Catalog-aware metadata reads and writes with explicit namespace (catName, dbName, fullTableName) |
supported | Routed to the resolved catalog/backend normally. |
Legacy reads and writes using catalog<separator>db database names |
supported | Routed by externalized database name; table/database objects are rewritten on the way back. |
Apache 3.1.3 wrapper RPCs against older Hortonworks 3.1.0.x backends |
degraded | Proxy retries selected *_req APIs through older legacy methods such as get_table. |
Apache 3.1.3 positional RPCs against Hive 4.1.x backends (catalog.<name>.runtime-profile=APACHE_4_1_0) |
degraded | Backend adapter auto-upgrades the two positional methods Hive 4 removed (get_table, get_table_objects_by_name) to their *_req equivalents; everything else flows through binary-compatible Thrift delegation unchanged. |
Hive 4.1.x clients (compatibility.frontend-profile=APACHE_4_1_0) against Apache 3.1.3 or Hortonworks 3.1.0.x backends |
degraded | Frontend bridge downgrades selected Hive 4-only *_req wrappers on read and standard-DDL paths (get_database_req, get_table_req, get_partition*_req, create_table_req, alter_table_req, etc.) to their positional Apache 3.1.3 equivalents; the 199 shared methods flow through binary-compatible delegation. |
Hive 4-only APIs (data connectors, scheduled queries, stored procedures, packages, ACID v2 extensions) on APACHE_4_1_0 front door |
rejected | Responded with TApplicationException UNKNOWN_METHOD; there is no safe downgrade to an Apache 3.1.3 backend for these RPCs. |
| Session-level and global read-only RPCs without catalog context | degraded | Routed to routing.default-catalog, including getMetaConf, get_all_functions, get_metastore_db_uuid, get_current_notificationEventId, get_open_txns, and get_open_txns_info. |
Read-only service APIs missing on a backend (TApplicationException UNKNOWN_METHOD on notifications, privilege refresh/introspection, token/key listings except delegation-token issuance, txn/lock/compaction status) |
degraded | Proxy returns an empty compatibility response instead of failing the caller. A backend that has the method but fails the call (any other TApplicationException type, transport failure, or MetaException) is reported as an error, so an empty answer never masquerades as real ACID, lock, or privilege state. |
Optional service reads (get_active_resource_plan, get_all_resource_plans, get_runtime_stats) |
degraded | Proxy returns an empty compatibility response on any backend failure, because these RPCs only expose optional workload-management and diagnostic state. |
ACID / txn / lock lifecycle RPCs without routable namespace (open_txns, commit_txn, abort_txn, check_lock, unlock, heartbeat) |
degraded | Pinned to routing.default-catalog; non-ACID SELECT locks, eligible non-transactional NO_TXN DDL locks and non-transactional write locks on non-default catalogs can use the proxy's synthetic lock shim, but this is still not a distributed ACID coordinator and the shim does not serialize concurrent writers. |
| Global write operations without clear catalog context | rejected | Proxy enforces deterministic routing and explicit namespace ownership: namespace-less service writes such as setMetaConf, grant_role, revoke_role, and add_token fail safely instead of guessing a catalog. |
Catalog registry management (create_catalog, drop_catalog) |
rejected | Catalog ownership is policy-managed in proxy config, not delegated to client RPCs. |
| HDP-only front-door methods with an explicit Apache bridge mapping | supported | Proxy adapts selected Hortonworks request-wrapper methods to Apache equivalents. |
HDP-only methods that require a matching Hortonworks runtime (add_write_notification_log, get_tables_ext, get_all_materialized_view_objects_for_rewriting) |
passthrough | Forwarded only to compatible Hortonworks backends/front doors when configured. |
| HDP-only methods without a safe Apache mapping | rejected | Proxy fails explicitly instead of returning a misleading success response. |
This matrix is the public contract for how to think about the proxy. It is intentionally grouped by client shape, advertised front door, backend runtime, auth mode, and method family rather than by individual thrift method names. The table is generated from capabilities.yaml and each capability row is linked to smoke-test coverage in the test suite.
For a method-level spreadsheet view of backend support, routing mode, fallback strategy, and semantic risk, see COMPATIBILITY.md.
Refresh the generated table with:
mvn -o -q -Dtest=CapabilityMatrixDocSyncTest -Dcapabilities.updateReadme=true test| Client version | Front-door profile | Backend profile | Auth mode | Method families | Expected result |
|---|---|---|---|---|---|
Apache Hive / Spark clients that speak Apache HMS 3.1.3 request wrappers |
APACHE_3_1_3 |
APACHE_3_1_3 |
NONE or KERBEROS |
catalog-aware reads/writes, legacy catalog<separator>db routing, view rewrite |
Supported as the baseline deployment. |
Apache Hive / Spark clients that speak Apache HMS 3.1.3 request wrappers |
APACHE_3_1_3 |
Hortonworks 3.1.0.x |
NONE or KERBEROS |
read paths, selected metadata writes that can fall back from *_req APIs |
Supported with compatibility downgrade; some calls are degraded to legacy RPCs. |
Hive 4.1.x clients (beeline, JDBC, Spark, Trino-Hive) that speak the Hive 4 request-wrapper Thrift API |
APACHE_4_1_0 |
APACHE_3_1_3, Hortonworks 3.1.0.x, or APACHE_4_1_0 |
NONE or KERBEROS |
199 read/write methods shared with Apache 3.1.3 (binary-compatible), selected Hive 4-only *_req wrappers downgraded to positional Apache 3.1.3 RPCs, lock with the Hive 4-only EXCL_WRITE lock type downgraded to EXCLUSIVE |
Supported with compatibility downgrade for read + standard DDL; Hive 4-only APIs (data connectors, scheduled queries, stored procedures, packages, ACID v2 extensions) respond with TApplicationException UNKNOWN_METHOD. |
Apache Hive / Spark clients that speak Apache HMS 3.1.3 request wrappers |
APACHE_3_1_3 |
Hive 4.1.x |
NONE or KERBEROS |
199 read/write methods shared with Hive 4 via binary-compatible Thrift delegation, positional get_table and get_table_objects_by_name (removed in Hive 4) auto-upgraded to their *_req equivalents |
Supported with compatibility upgrade for the two positional methods Hive 4 removed; all other RPCs flow through unchanged. Hive 4-only APIs (data connectors, scheduled queries, etc.) are not reachable from an Apache 3.1.3 front door because Apache 3.1.3 has no Thrift bindings for them. |
Hortonworks clients that only expect a Hortonworks front-door identity via getVersion() |
HORTONWORKS_* without standalone jar |
APACHE_3_1_3 or Hortonworks 3.1.0.x |
NONE or KERBEROS |
overlapping Apache/HDP method families | Supported when changing the advertised profile is enough for the client. |
| Hortonworks clients that call HDP-only thrift request-wrapper methods | HORTONWORKS_* with standalone jar |
Hortonworks 3.1.0.x |
NONE or KERBEROS |
mapped HDP-only methods, runtime-specific passthrough methods | Supported when the matching Hortonworks front-door and backend runtime jars are configured. |
| Hortonworks clients that call HDP-only thrift request-wrapper methods | HORTONWORKS_* with standalone jar |
APACHE_3_1_3 |
NONE or KERBEROS |
HDP-only passthrough methods such as add_write_notification_log |
Rejected explicitly when the target backend does not provide a compatible Hortonworks runtime. |
| HiveServer2 / Beeline SQL workloads across multiple catalogs | APACHE_3_1_3 or HORTONWORKS_* |
mixed Apache + Hortonworks backends | NONE or KERBEROS |
reads, DDL/DML, namespace rewrite, optional view rewrite | Supported as long as routing can resolve the target catalog. |
| HiveServer2 / direct HMS clients using txn/lock lifecycle RPCs without namespace in the payload | any | mixed Apache + Hortonworks backends | NONE or KERBEROS |
open_txns, commit_txn, abort_txn, check_lock, unlock, heartbeat |
Degraded: pinned to routing.default-catalog; eligible non-ACID SELECT, NO_TXN DDL and non-transactional write (INSERT/UPDATE/DELETE) locks can still be synthesized on non-default catalogs, but otherwise treat this as a single-catalog control plane unless you validated otherwise. |
| Kerberized HiveServer2 / HMS clients that require end-user identity on the backend | any | any | KERBEROS with optional impersonation |
front-door SASL, local delegation-token issuance, backend set_ugi() impersonation |
Supported when proxy-user rules and backend impersonation permissions are configured correctly. |
| Clients attempting mutations without explicit namespace ownership or dynamic catalog registry management | any | any | NONE or KERBEROS |
policy-guarded ambiguous mutations, create_catalog, drop_catalog |
Safely failed by design to preserve deterministic routing, explicit namespace ownership, and no silent split-brain writes. |
- legacy database references without a catalog prefix are routed to
routing.default-catalog - this keeps Spark/Hive compatibility while still preserving deterministic routing and avoiding silent split-brain metadata writes
- in practice this means ACID write lifecycle is supported only for the default catalog unless the
request payload itself carries routable namespace information such as
dbNameorfullTableName - the proxy can synthesize non-ACID
SHARED_READSELECTlocks, eligible non-transactionalNO_TXNDDL locks and non-transactional write locks (INSERT,UPDATE,DELETE) for non-default catalogs so HiveServer2 read, non-ACID DDL and non-ACIDINSERTpaths stop depending on backend txn state alignment - that synthetic lock state is in-memory by default and can be moved to ZooKeeper for multi-instance proxy failover, but it still does not make ACID writes or write-id coordination multi-catalog safe
- Hive ACID, locks, tokens, and other truly global metastore operations still need careful validation in your environment before turning them on behind a multi-catalog proxy
The proxy can also apply optional latency-aware backend handling for slow or intermittently failing metastores. This layer is disabled by default, so existing deployments keep the previous routing behavior until you opt in.
When enabled, the proxy can:
- keep a per-catalog latency budget via
catalog.<name>.latency-budget-ms - track per-backend latency EWMA and derive adaptive backend socket timeouts from recent response times
- open a circuit after repeated transport or time-budget failures, then allow a half-open retry after
routing.circuit-breaker.open-state-ms - poll backend readiness in the background instead of only probing on demand through
/readyz - run only safe read-only fanout RPCs in parallel when
routing.hedged-read.enabled=true - omit degraded backends from those safe fanout reads when
routing.degraded-routing-policy=SAFE_FANOUT_READS
This is intentionally narrow in scope: hedged reads and degraded omission apply only to safe
read-only fanout methods, currently get_all_databases, get_databases, and get_table_meta.
Single-backend writes and namespace-sensitive mutations still follow the deterministic routing model
described above and do not race multiple metastores.
Non-impersonated calls to a catalog go through a borrow/return pool of backend Thrift sessions sized
by catalog.<name>.shared-session-pool-size. The default is 1, which preserves the previous
single-session, serialized behavior — concurrent shared calls will queue on a single Thrift
transport. To benefit from parallelism, set this explicitly per catalog, e.g.:
catalog.catalog1.shared-session-pool-size=8Trade-offs to consider when raising it:
- More idle Thrift sessions held open against the backend HMS (with proportional Kerberos cost when
KERBEROSis enabled on the backend). reconnectShareddrains the full pool, so larger pools take longer to swap on a runtime profile change or forced reconnect.
When impersonation is enabled on the catalog, real authenticated client traffic does not use this
pool — each caller is served from the per-user impersonation client cache governed by
catalog.<name>.impersonation-max-clients and catalog.<name>.impersonation-client-idle-ttl-ms.
Each cached user owns an internal pool of backend Thrift sessions sized by
catalog.<name>.impersonation-pool-max-size (default 4); idle sessions inside that pool can be
closed via catalog.<name>.impersonation-session-idle-ttl-ms (default 0 = never). Concurrent
calls from the same user (or the HiveServer2 driving them) borrow distinct sessions in parallel
instead of serializing on a single Thrift transport. Grafana surfaces this as
hms_proxy_impersonation_pool_users, hms_proxy_impersonation_pool_sessions{state=active|idle},
hms_proxy_impersonation_session_acquire_timeouts_total (per-user borrow timeouts) and
hms_proxy_impersonation_session_evictions_total{reason=idle|transport_failure|user_evicted|user_capacity}.
The shared pool then carries only:
- internal proxy-driven calls such as backend health probes and reconnect bookkeeping;
- requests authenticated as a service principal listed under
security.service-principals.*, which by design bypass impersonation and reuse a shared backend session.
In a pure-impersonation deployment with no service-principal client traffic, raising
shared-session-pool-size therefore has no effect on application throughput. Tune it when at least
one catalog has impersonation-enabled=false, or when service-principal callers (for example
HiveServer2 itself) drive a meaningful share of requests against the catalog.
mvn -o test
mvn -o package
mvn -q -DforceStdout help:evaluate -Dexpression=project.versionBuild version is computed from git for every commit in the form 0.1.<git-distance>-<short-sha>.
GitHub Actions publishes prerelease builds automatically:
- every push to
maincreates abuild-<project.version>tag and attaches the built jars to a prerelease - a nightly run is scheduled for
00:00 UTCevery day and publishes anightly-YYYYMMDDprerelease from the currentmainhead - each prerelease also gets auto-generated GitHub release notes for that tag
Manual releases are published through the Release workflow:
- start it with
workflow_dispatch - pass only
major_minorsuch as1.7 - the workflow computes the next patch version automatically and publishes a GitHub release tag like
v1.7.0,v1.7.1, and so on
java -jar "target/hms-proxy-$(mvn -q -DforceStdout help:evaluate -Dexpression=project.version)-fat.jar" /etc/hms-proxy/hms-proxy.propertiesmvn package produces both a regular jar and a runnable fat jar with classifier fat.
The fat jar file name changes with every new commit.
For Java 17+ with Hadoop 2.x Kerberos libraries, start with:
java \
--add-opens=java.security.jgss/sun.security.krb5=ALL-UNNAMED \
--add-exports=java.security.jgss/sun.security.krb5=ALL-UNNAMED \
-Djava.security.krb5.conf=/etc/krb5.conf \
-jar "target/hms-proxy-$(mvn -q -DforceStdout help:evaluate -Dexpression=project.version)-fat.jar" /etc/hms-proxy/hms-proxy.propertieslibthrift accepts client sockets with an infinite read timeout, so a client that disappears
without FIN/RST — a network failure, a process killed behind NAT — leaves its worker thread
blocked in read forever. The worker pool is bounded by server.max-worker-threads, so the
listener slowly drains until it stops accepting new connections. Every listener therefore
applies a bounded socket lifetime:
| Key | Default | Meaning |
|---|---|---|
server.client-socket-timeout-ms |
600000 |
Read timeout on an accepted connection; 0 restores the unbounded libthrift behavior |
server.tcp-keepalive |
true |
SO_KEEPALIVE on accepted connections |
server.tcp-keepalive-idle-seconds |
120 |
Idle time before the first keepalive probe |
server.tcp-keepalive-interval-seconds |
30 |
Interval between keepalive probes |
server.tcp-keepalive-count |
4 |
Failed probes before the connection is dropped |
server.shutdown-timeout-seconds |
30 |
Bound on the ordered teardown after SIGTERM |
client-socket-timeout-ms is a read timeout: it bounds how long a worker waits for the next
request on an established connection and does not abort long server-side calls such as
drop_table with purge, because no read is in flight while a request is being processed. The
default is the order of magnitude of Hive's own hive.metastore.client.socket.timeout (600s);
clients wrapped in RetryingMetaStoreClient (HiveServer2, Spark) reconnect transparently when an
idle connection is recycled.
The keepalive timers bound dead-peer detection at idle + interval * count seconds — 4 minutes
with the defaults, instead of the roughly two hours the OS-wide tcp_keepalive_* defaults give.
Tuning the timers needs per-socket keepalive options (Linux and macOS); on a platform without
them the proxy logs once per listener and falls back to the OS-wide timers, keeping plain
SO_KEEPALIVE.
On SIGTERM the shutdown hook stops the primary listener and then waits for the main thread to
close, in order, the additional frontend listeners, the management listener, the router backends
and the front-door security. The JVM halts as soon as the last shutdown hook returns, so this
wait is what keeps in-flight requests on the additional frontends from being cut off and backend
resources from being left open. It is bounded by server.shutdown-timeout-seconds; when the
teardown does not finish in time the proxy logs a warning and lets the JVM halt anyway.
The proxy validates its properties file at startup and refuses to start on a value it cannot interpret. A misconfigured proxy that runs and quietly does the opposite of what the file says is worse than one that does not start.
Booleans accept only true and false, in any case. yes, on, 1 and typos such as ture
are startup errors naming the key and the value. Previously they were read as false, which
silently disabled impersonation, the management listener, or hedged reads.
Enum values — every mode, profile and policy key — are case-insensitive, so
access-mode=read_only and frontend-profile=apache_3_1_3 are accepted everywhere. An unknown
value reports the key and the accepted constants.
Durations in HiveConf keys such as hive.metastore.client.socket.timeout follow Hive: an
integer with an optional unit suffix ns, us, ms, s/sec, m/min, h/hour, d/day
(long spellings such as 600sec, 5min work too). A bare number means seconds. A value Hive
itself would reject, such as 1.5s, is logged as a WARN naming the value and the applied
default.
Contradictory combinations are startup errors rather than silent behavior swaps:
| Combination | Result |
|---|---|
catalog.<name>.write-db-whitelist without access-mode=READ_WRITE_DB_WHITELIST |
error: the whitelist would be ignored and writes would stay open to every database |
access-mode=READ_WRITE_DB_WHITELIST with an empty whitelist |
error: use READ_ONLY to forbid all writes |
synthetic-read-lock.store.mode=IN_MEMORY with any synthetic-read-lock.store.zookeeper.* key |
error: locks would silently stay in memory |
Two listeners on the same host:port |
error, including a wildcard bind host: 0.0.0.0:9083 conflicts with 127.0.0.1:9083 |
Listener conflict checks cover the primary listener, the management listener (including its
default server.port + 1000) and every additional-frontends.<name> entry. Host names are
compared as configured, without DNS resolution, so localhost versus 127.0.0.1 still surfaces
at bind time rather than at validation time.
The proxy can expose a lightweight HTTP listener for health checks, readiness, and Prometheus
metrics. The listener is disabled by default and turns on automatically when management.port
is configured:
management.bind-host=0.0.0.0
management.port=19083You can also enable it explicitly:
management.enabled=true
# Optional; defaults to server.bind-host
management.bind-host=0.0.0.0
# Optional; defaults to server.port + 1000
management.port=19083
# Optional; handler threads for the management listener, defaults to 4
management.threads=4
# Optional; how long readiness probe results are reused, in ms; defaults to 2000, 0 disables caching
management.readiness-cache-ms=2000Quick checks:
curl -s http://127.0.0.1:19083/healthz
curl -s http://127.0.0.1:19083/readyz
curl -s http://127.0.0.1:19083/metricsThe listener serves requests from a small dedicated thread pool (management.threads), so a slow
/readyz never blocks /healthz or /metrics. Keep at least two threads.
The management endpoints have no authentication or authorization. /readyz exposes Kerberos
principals and per-backend lastError strings that usually contain internal hostnames. Bind
management.bind-host to an interface reachable only from your monitoring and orchestration
network, or otherwise restrict the port with firewall rules or a network policy. Do not publish it
alongside the client-facing Thrift port.
Available endpoints:
/healthzreturns process liveness and uptime/readyzchecks backend connectivity, returns per-backendconnected/degradedstate, and includes Kerberos login status plus TGT freshness for front-door and outbound backend credentials/metricsexposes Prometheus text format metrics
/healthz is intended for simple liveness checks and only answers whether the proxy process is up.
/readyz is intended for load balancers, orchestration probes, and operational diagnostics. The
response includes:
- overall readiness status
probeAgeMs, the age of the probe results the response was built from- backend connectivity summary
- per-backend
connected,degraded,lastSuccessEpochSecond,lastFailureEpochSecond,lastProbeEpochSecond,lastLatencyMs,latencyEwmaMs,baselineTimeoutMs,adaptiveTimeoutMs,latencyBudgetMs,circuitState,consecutiveFailures,circuitRetryAtEpochMs, andlastError - Kerberos status for the front door and outbound backend credentials, including
state(DISABLED,PENDING,ACTIVE,STALE,FAILED),loggedIn, andhealthy - Kerberos TGT freshness via
tgtExpiresAtEpochSecondandsecondsUntilExpirywhen available
Kerberos status is read from the logins the proxy already uses. The Kerberos probe never reinstalls
the process-wide security configuration and never performs a keytab login of its own, so it does not
race with live SASL handshakes and adds no KDC traffic. Note that the on-demand backend connectivity
probes above still open a backend session per scrape, which does perform a keytab login when the
backend uses SASL; enable routing.backend-state-polling.enabled=true to move that off the request
path.
Readiness fails only when the front-door login user is missing or has no Kerberos credentials at
all. An aged-out TGT is reported as state=STALE with secondsUntilExpiry<=0 instead of failing
readiness, because Hadoop does not renew the TGT of a keytab login on its own while the SASL
acceptor keeps authenticating clients with the keytab service keys. Backend Kerberos is diagnostic
only: PENDING until the first backend session opens, STALE when the last recorded backend login
is unusable - every backend session performs its own keytab login, so real backend health comes from
the connectivity probes. Because the probe no longer logs in, a rotated or revoked keytab is not
detected by /readyz itself; it surfaces through failing client authentication or backend
connectivity.
If routing.backend-state-polling.enabled=true, readiness reflects the most recent background probe
results. Otherwise /readyz measures backend probe latency on demand and returns the same fields.
Backend and Kerberos probes are the expensive part of /readyz, so their results are reused for
management.readiness-cache-ms (2s by default) and refreshed by at most one request at a time:
concurrent callers are served the previous result rather than fanning out more network checks. The
per-backend state fields are always rendered from current in-memory runtime state, and probeAgeMs
tells you how stale the underlying probe data is. Set management.readiness-cache-ms=0 to probe on
every request.
Current Prometheus metrics:
hms_proxy_requests_total{method,catalog,backend,status}hms_proxy_request_duration_seconds{method,catalog,backend}hms_proxy_backend_failures_total{backend,exception}hms_proxy_backend_fallback_total{method,from_api,to_api}hms_proxy_routing_ambiguous_totalhms_proxy_default_catalog_routed_total{method}hms_proxy_lock_request_split_total{catalog}hms_proxy_filtered_objects_total{method,catalog,object_type}hms_proxy_synthetic_read_lock_events_total{operation,catalog,store_mode,result}hms_proxy_synthetic_read_lock_store_failures_total{operation,store_mode,exception}hms_proxy_synthetic_read_lock_handoffs_total{operation,catalog,store_mode}hms_proxy_synthetic_read_locks_active{store_mode}hms_proxy_synthetic_read_lock_store_info{store_mode}hms_proxy_backend_session_acquire_timeouts_total{catalog,operation}hms_proxy_adaptive_timeout_reconnect_total{catalog}hms_proxy_adaptive_timeout_reconnect_skipped_total{catalog,reason}hms_proxy_iceberg_pointer_guard_events_total{catalog,outcome}hms_proxy_rest_requests_total{prefix,route,status}hms_proxy_rest_request_duration_seconds{prefix,route}hms_proxy_rest_listener_info{bind_host,port}
Example Prometheus scrape config:
scrape_configs:
- job_name: hms-proxy
static_configs:
- targets:
- hms-proxy-01.example.com:19083Metric semantics:
statusinhms_proxy_requests_totalis one ofok,error,fallback, orthrottledcatalog=all, backend=fanoutmeans a request was broadcast to multiple backendshms_proxy_backend_failures_totalcounts backend-side invocation failures grouped by backend and exception typehms_proxy_backend_fallback_totalcounts compatibility fallbacks returned after backend failureshms_proxy_routing_ambiguous_totalcounts requests that safely failed because the proxy saw conflicting namespace hints and refused to guesshms_proxy_default_catalog_routed_totalcounts requests that were routed to the default catalog because no explicit catalog namespace was presenthms_proxy_lock_request_split_totalcounts lock requests that named several catalogs and were routed to one of them, the other components dropped from the request that reached the backendhms_proxy_rate_limited_totalcounts requests rejected by overload protection with labels for the limiting dimension, configured scope, method family, and resolved cataloghms_proxy_filtered_objects_totalcounts databases or tables hidden by selective federation exposure rules before they are returned to the clienthms_proxy_synthetic_read_lock_events_totaltracks synthetic lock shim lifecycle transitions such asacquire,check_lock,heartbeat,unlock,release_txn, andcleanuphms_proxy_synthetic_read_lock_store_failures_totalcounts in-memory or ZooKeeper store failures grouped by operation and exception typehms_proxy_synthetic_read_lock_handoffs_totalcounts cases where one proxy instance continues serving a synthetic lock originally acquired through another instancehms_proxy_synthetic_read_locks_activeexposes the number of currently visible synthetic locks for the configured store backend; it is adjusted in place on every acquire/release and re-synchronized with the store by the background expiry sweep (every 30s), so lock operations never pay for a store-wide listinghms_proxy_synthetic_read_lock_store_infois a constant-info gauge that marks whether this proxy runs within_memoryorzookeepersynthetic lock storagehms_proxy_backend_session_acquire_timeouts_totalcounts fail-fast events when the shared backend metastore session pool runs out of permits within the catalog'slatencyBudgetMs(or 30s default);operation=borrowcovers regular RPC dispatch,operation=reconnectcovers admin reconnect attempts that could not quiesce the poolhms_proxy_adaptive_timeout_reconnect_totalcounts how often the adaptive socket timeout reconnected the shared backend client (and forced impersonation-cache eviction); use it to spot reconnect storms under volatile latencyhms_proxy_adaptive_timeout_reconnect_skipped_totalcounts adaptive-timeout adjustments suppressed by the throttles (reason=hysteresisfor sub-threshold deltas,reason=cooldownfor events too close to a previous reconnect)hms_proxy_iceberg_pointer_guard_events_totalcounts the Iceberg pointer guard's decisions onalter_table:repaired(the request would have erased the record's Iceberg state and was merged over it),forward_commit(an Iceberg commit built on the current pointer, passed through),not_iceberg(the record holds no pointer, remembered as such),cache_suppressed(answered from that memory, no backend read),read_failed(the record could not be read and the alter was passed through). Everything exceptcache_suppressedcost oneget_table, so the added round trips and the cache hit rate are both readable from this one metrichms_proxy_rest_requests_totalcounts Iceberg REST HTTP requests by catalog prefix, route, and terminal HTTP statushms_proxy_rest_request_duration_secondsmeasures Iceberg REST request duration grouped by catalog prefix and routehms_proxy_rest_listener_infois a constant-info gauge that exposes the configured bind host and port of the Iceberg REST listener
Despite the historical synthetic_read_lock metric names, the shim now also serves eligible
non-transactional NO_TXN DDL locks and non-transactional write locks on non-default catalogs.
The bundled Grafana dashboard at monitoring/grafana/hms-proxy-dashboard.json includes panels for
synthetic lock activity, handoffs, store failures, and active lock counts, plus a store_mode
templating variable for quickly switching between in_memory and zookeeper views.
The proxy emits one structured audit log record per request through the logger
io.github.mmalykhin.hmsproxy.audit. Each record is a single-line JSON object with fields such as
requestId, method, operationClass, catalog, backend, status, durationMs,
remoteAddress, authenticatedUser, routed, fanout, fallback, and
defaultCatalogRouted.
Client-controlled values (authenticatedUser, remoteAddress) are JSON-escaped, including every
control character below 0x20, so a hostile principal or address cannot break the record for log
collectors. Non-ASCII characters are emitted verbatim as UTF-8.
Example:
{"event":"hms_proxy_audit","requestId":42,"method":"get_table","operationClass":"metadata_read","catalog":"catalog1","backend":"catalog1","status":"ok","durationMs":8,"routed":true,"fanout":false,"fallback":false,"defaultCatalogRouted":false,"remoteAddress":"10.20.30.40","authenticatedUser":"alice@EXAMPLE.COM"}The bundled log4j.properties routes this logger to its own appender, logs/hms-proxy-audit.log
(rolling at 100MB with 10 backups), with no layout prefix so the file stays valid JSON lines, and
with additivity off so audit records do not mix into the general log. Point it somewhere else, or
ship it to your log pipeline, by overriding the auditFile appender in your own
log4j.properties. Silencing the logger below INFO also skips building the record, so nothing is
computed for output nobody reads.
A ready-to-import Grafana dashboard is included in
monitoring/grafana/hms-proxy-dashboard.json. It covers request rate, latency, backend failures,
fallbacks, default-catalog routing, and ambiguous routing events, plus an Iceberg REST row:
request rate, error ratio and latency quantiles of the REST listener, breakdowns by HTTP status,
catalog prefix and route, and a listener-up stat.
You can publish only part of a backend namespace while keeping routing and write policy unchanged. This is useful for gradual migration, safer multi-tenant rollout, and avoiding accidental metadata exposure.
Per catalog:
# Backward-compatible default
catalog.catalog1.expose-mode=ALLOW_ALL
# Safer rollout mode: only matched objects are visible
catalog.catalog1.expose-mode=DENY_BY_DEFAULT
catalog.catalog1.expose-db-patterns=sales,finance_.*
catalog.catalog1.expose-table-patterns.sales=orders_.*,events
catalog.catalog1.expose-table-patterns.finance_.*=audit_.*Rules:
- regexes are matched case-insensitively
- matching is done against backend names, not externalized proxy names
- matching uses the whole string;
salesmatches onlysales, whilesales_.*matchessales_eu catalog.<name>.expose-db-patternsis an allowlist for backend database names inside that catalogcatalog.<name>.expose-table-patterns.<dbRegex>is an allowlist for table names inside databases whose backend db name matches<dbRegex>- table rules narrow visibility for matching databases; unmatched tables are filtered out
- with
DENY_BY_DEFAULT, databases without a matching db rule or table rule are hidden - a matching table rule can expose a database even when no db-level rule is present, but only matching tables inside that database stay visible
The filter is applied on metadata read paths such as get_all_databases, get_databases,
get_table*, get_tables*, get_table_meta, and Hortonworks get_tables_ext.
Behavior by API shape:
- list-style RPCs such as
get_all_databases,get_all_tables,get_tables,get_table_meta, andget_tables_extsilently drop hidden objects from the response - direct lookups such as
get_database,get_table, andget_table_reqfail as "not found" when the target object is filtered out hms_proxy_filtered_objects_total{method,catalog,object_type}counts both hidden databases and hidden tables
Examples:
# 1. Publish only one database during migration
catalog.catalog1.expose-mode=DENY_BY_DEFAULT
catalog.catalog1.expose-db-patterns=sales
# 2. Publish one database, but only selected tables inside it
catalog.catalog1.expose-mode=DENY_BY_DEFAULT
catalog.catalog1.expose-table-patterns.sales=orders_.*,customers
# 3. Keep the catalog open by default, but narrow one sensitive database
catalog.catalog1.expose-mode=ALLOW_ALL
catalog.catalog1.expose-table-patterns.audit=.*_publicWhat those examples mean:
- example 1 exposes only backend db
sales - example 2 makes backend db
salesvisible and exposes only tables matchingorders_.*pluscustomers - example 3 leaves all databases visible, but in backend db
auditreturns only tables matching.*_public
For non-default catalogs, remember that filters still use backend db names such as sales, not
external names like catalog2__sales.
You can ask the proxy to protect table creation or table alteration when the incoming table metadata marks a managed table as transactional:
transactional=true- any non-empty
transactional_properties
The rule applies to every create_table* and alter_table* RPC, including
create_table_with_environment_context (the call HiveMetaStoreClient actually sends for
createTable), create_table_with_constraints, and alter_table_with_cascade. The *_req
wrappers from Hive 4 and Hortonworks front doors are unwrapped into these RPCs before the guard
runs, so they are covered too. It is evaluated only for MANAGED_TABLE. External tables are left
unchanged.
Reject mode:
guard.transactional-ddl.mode=REJECT_TRANSACTIONALRewrite mode:
guard.transactional-ddl.mode=REWRITE_TRANSACTIONAL_TO_EXTERNALIn rewrite mode the proxy rewrites the incoming table to EXTERNAL_TABLE, adds
external.table.purge=true, and removes transactional plus transactional_properties.
If you also need physical file deletion on Apache 3.1.3 backends, combine this with
federation.external-table-drop-purge.mode=BEST_EFFORT and a per-catalog allowlist via
catalog.<name>.conf.hms.proxy.external-table-drop-purge.allowed-prefixes.
You can also scope it to specific client IPs or CIDR ranges:
guard.transactional-ddl.mode=REJECT_TRANSACTIONAL
guard.transactional-ddl.client-addresses=10.10.0.15,10.20.0.0/16,2001:db8::/64If guard.transactional-ddl.client-addresses is set, the rule is evaluated only for matching
clients. If it is omitted, all clients are covered.
The proxy can also reject bursts before they become backend overload. This protection is local to each proxy instance and uses token-bucket limits with:
- a steady refill rate via
requests-per-second - an optional short burst allowance via
burst - independent buckets, so a request must pass every configured scope that applies to it
Supported scopes:
- per authenticated client principal:
rate-limit.principal.* - per exact source IP:
rate-limit.source.* - per source CIDR rule:
rate-limit.source-cidr.<name>.* - per HMS method family:
rate-limit.method-family.<family>.* - per logical catalog:
rate-limit.catalog.<catalog>.* - per high-risk RPC class:
rate-limit.rpc-class.<class>.*
Supported method families:
metadata_readmetadata_writeservice_global_readservice_global_writeacid_namespace_bound_writeacid_id_bound_lifecycleadmin_introspectioncompatibility_only_rpc
Supported RPC classes:
writeddltxnlock
Important behavior:
- per-principal limits apply only when the front door exposes an authenticated user, for example Kerberos/SASL
rate-limit.source-cidr.<name>is an aggregate bucket for the whole configured CIDR rule, not a separate bucket per IP- if one source IP matches several CIDR rules, all matching rules are enforced
- per-catalog limits are enforced when the request actually resolves or touches a catalog/backend; fanout reads can therefore consume more than one catalog bucket
- one RPC can match several classes at once, for example a lock/txn write can count toward
write,txn, andlock - when a limit is exceeded the client receives a
MetaExceptionand the request is recorded withstatus="throttled"
Example:
# Per authenticated principal
rate-limit.principal.requests-per-second=60
rate-limit.principal.burst=120
# Per exact source IP
rate-limit.source.requests-per-second=30
rate-limit.source.burst=60
# Aggregate bucket for a whole subnet or pool
rate-limit.source-cidr.hs2-pool.cidrs=10.10.0.0/16,10.20.0.0/16
rate-limit.source-cidr.hs2-pool.requests-per-second=200
rate-limit.source-cidr.hs2-pool.burst=300
# Method-family shaping
rate-limit.method-family.metadata_read.requests-per-second=600
rate-limit.method-family.metadata_read.burst=1000
# Per catalog
rate-limit.catalog.catalog1.requests-per-second=300
rate-limit.catalog.catalog2.requests-per-second=120
# Extra protection for high-risk RPC classes
rate-limit.rpc-class.write.requests-per-second=80
rate-limit.rpc-class.ddl.requests-per-second=15
rate-limit.rpc-class.txn.requests-per-second=30
rate-limit.rpc-class.lock.requests-per-second=50Recommended production starting point:
- use
principalfor runaway HS2 sessions or bad end users - use
sourceandsource-cidrfor tooling, scanners, or large client pools - use
method-family.metadata_readto cap scan-heavy metadata discovery - use
catalog.<name>to keep one hot catalog from starving the others - keep stricter limits on
ddl,txn, andlockthan on plain metadata reads
You can also choose which HMS version the proxy advertises on the front door:
compatibility.frontend-profile=APACHE_3_1_3or for Hortonworks clients:
compatibility.frontend-profile=HORTONWORKS_3_1_0_3_1_0_78or for HDP 3.1.0.3.1.5.6150-1 clients:
compatibility.frontend-profile=HORTONWORKS_3_1_0_3_1_5_6150_1That changes the value returned by getVersion() while keeping the same proxy routing logic.
The Thrift protocol has no version-negotiation handshake, so the proxy cannot tell which Hive
version a client speaks from the first request (much like HiveServer2's split between binary
and HTTP protocols). When two different client populations need different getVersion()
identities or different runtime jars, run multiple Thrift listeners on different ports:
server.port=9083
compatibility.frontend-profile=APACHE_3_1_3
additional-frontends=hdp
additional-frontends.hdp.port=9084
additional-frontends.hdp.frontend-profile=HORTONWORKS_3_1_0_3_1_5_6150_1
additional-frontends.hdp.standalone-metastore-jar=hive-metastore/hive-standalone-metastore-3.1.0.3.1.5.6150-1.jarThe primary listener stays on server.port with its own compatibility.frontend-profile.
Each additional-frontends.<name>.* block defines an independent listener:
| Key | Required | Default |
|---|---|---|
additional-frontends.<name>.port |
yes | — |
additional-frontends.<name>.frontend-profile |
yes | — |
additional-frontends.<name>.standalone-metastore-jar |
yes for non-APACHE_3_1_3 |
— |
additional-frontends.<name>.bind-host |
no | server.bind-host |
additional-frontends.<name>.min-worker-threads |
no | server.min-worker-threads |
additional-frontends.<name>.max-worker-threads |
no | server.max-worker-threads |
additional-frontends.<name>.client-socket-timeout-ms |
no | server.client-socket-timeout-ms |
additional-frontends.<name>.tcp-keepalive |
no | server.tcp-keepalive |
additional-frontends.<name>.tcp-keepalive-idle-seconds |
no | server.tcp-keepalive-idle-seconds |
additional-frontends.<name>.tcp-keepalive-interval-seconds |
no | server.tcp-keepalive-interval-seconds |
additional-frontends.<name>.tcp-keepalive-count |
no | server.tcp-keepalive-count |
All listeners share the same RoutingMetaStoreProxy, federation, security (FrontDoorSecurity
including SASL/Kerberos), audit and Prometheus metrics. Only the wire-level Thrift API
differs per port.
Additional listeners run on daemon threads and are stopped before the router and the front-door security are closed, so a failure while starting one listener cannot leave the JVM alive with a port still held. See Front-door socket lifetime and shutdown.
Every listener must own its host:port. Startup fails when an additional frontend collides with
the primary listener, with another additional frontend, or with the management listener — including
the management default server.port + 1000 — and a wildcard bind host such as 0.0.0.0 counts as a
collision with any host on the same port.
For a real Hortonworks front door, point the proxy to an HDP standalone-metastore jar:
compatibility.frontend-profile=HORTONWORKS_3_1_0_3_1_0_78
compatibility.frontend-standalone-metastore-jar=/opt/hms-proxy/hive-metastore/hive-standalone-metastore-3.1.0.3.1.0.0-78.jarFor HDP 3.1.0.3.1.5.6150-1, point the proxy to the matching jar:
compatibility.frontend-profile=HORTONWORKS_3_1_0_3_1_5_6150_1
compatibility.frontend-standalone-metastore-jar=/opt/hms-proxy/hive-metastore/hive-standalone-metastore-3.1.0.3.1.5.6150-1.jarFor isolated Hortonworks backend runtimes, you can also point the proxy to a backend jar explicitly:
compatibility.backend-standalone-metastore-jar=/opt/hms-proxy/hive-metastore/hive-standalone-metastore-3.1.0.3.1.0.0-78.jaror:
compatibility.backend-standalone-metastore-jar=/opt/hms-proxy/hive-metastore/hive-standalone-metastore-3.1.0.3.1.5.6150-1.jarFor Hortonworks backend runtimes the proxy forces hive.metastore.uri.selection=SEQUENTIAL
inside the isolated metastore client. This avoids a known HDP 3.1.0.x client bug in the
random URI selection path when HiveMetaStoreClient resolves backend metastore URIs.
Backend runtime is configured explicitly per catalog. If catalog.<name>.runtime-profile is not
set, the proxy uses APACHE_3_1_3 for that catalog:
catalog.hdp.runtime-profile=HORTONWORKS_3_1_0_3_1_0_78
catalog.hdp.backend-standalone-metastore-jar=/opt/hms-proxy/hive-metastore/hive-standalone-metastore-3.1.0.3.1.0.0-78.jar
catalog.apache.runtime-profile=APACHE_3_1_3For HDP 3.1.0.3.1.5.6150-1, configure the catalog the same way:
catalog.hdp.runtime-profile=HORTONWORKS_3_1_0_3_1_5_6150_1
catalog.hdp.backend-standalone-metastore-jar=/opt/hms-proxy/hive-metastore/hive-standalone-metastore-3.1.0.3.1.5.6150-1.jarWith that jar present, the proxy can open the selected Hortonworks backend runtime in an isolated
classloader. Runtime choice is not autodetected from the backend server version; it follows the
configured catalog.<name>.runtime-profile.
For the front door, the proxy instantiates the Hortonworks thrift Processor in an isolated
classloader and bridges overlapping RPCs to the internal Apache 3.1.3 handler automatically.
Selected HDP-only methods are also adapted to Apache equivalents:
get_database_req->get_databasefor HDP3.1.0.3.1.5.6150-1create_table_req->create_table/create_table_with_environment_context/create_table_with_constraintstruncate_table_req->truncate_tablealter_table_req->alter_table/alter_table_with_environment_contextalter_partitions_req->alter_partitions/alter_partitions_with_environment_contextrename_partition_req->rename_partitionget_partitions_by_names_req->get_partitions_by_namesupdate_table_column_statistics_req->set_aggr_stats_forupdate_partition_column_statistics_req->set_aggr_stats_foradd_write_notification_log-> direct Hortonworks passthrough only to a Hortonworks backendget_tables_ext-> direct Hortonworks passthrough only to a Hortonworks3.1.0.3.1.5.6150-1backendget_all_materialized_view_objects_for_rewriting-> direct Hortonworks passthrough only to a Hortonworks3.1.0.3.1.5.6150-1backend viarouting.default-catalog
Some HDP-only methods still do not have a safe Apache mapping, so they remain unsupported and fail explicitly rather than returning a misleading success response.
View/materialized-view notes:
- the proxy can rewrite SQL text only with
federation.view-text-rewrite.mode=REWRITE - rewrite is intentionally parser-less and conservative: a lexical scanner tracks table positions
(
FROM,JOIN,INTO,TABLE,UPDATE) and rewrites only the database qualifier of a reference standing in one of them; it does not parse the full Hive SQL grammar - string literals,
--and/* */comments, numbers and backquoted identifiers are skipped, so their content is never rewritten; column qualifiers and table aliases (t.colinselect t.col from sales.orders t) are left alone even when they collide with a database name catalog.db.tablereferences keep their catalog qualifier: outbound rewrite collapses<backend catalog>.<db>.<table>into the external database name, and any other catalog prefix is left untouched instead of being silently retargeted- cross-catalog references in inbound SQL like
catalog2__dim.table_xare internalized for the backend, but outbound rewrite is only guaranteed for the current table namespace - whatever cannot be resolved unambiguously stays untouched and is logged at
DEBUGbyViewDefinitionCompatibility; an unrewritten reference surfaces as an explicit backend error rather than as a silently corrupted view definition - by default only
viewExpandedTextis rewritten (federation.view-text-rewrite.preserve-original-text=true), so the client-facingviewOriginalTextis never mutated; set it tofalseif you also want the stored original SQL translated - dialect-specific text (variable substitution such as
${hiveconf:db}, macros, engine-specific hints) is out of scope and still needs validation in your environment
The bundled log4j.properties is the default configuration, and the proxy bootstraps it at runtime
if no appenders are configured, so a plain java -jar ... launch still gets logs. Defaults:
- root logger at
INFO, writing to stderr and tologs/hms-proxy.log(rolling at 50MB, 10 backups) - the proxy package
io.github.mmalykhin.hmsproxyatINFO - the audit logger
io.github.mmalykhin.hmsproxy.auditatINFOinto its ownlogs/hms-proxy-audit.log
Override any of it with your own log4j.properties on the classpath.
The binding is slf4j-reload4j. reload4j is a drop-in fork of the EOL log4j 1.2.17 with the known
CVEs removed, so the configuration format stays log4j 1.x and existing log4j.properties files keep
working unchanged. The pieces reload4j dropped — org.apache.log4j.jmx.*, the lf5 viewer and
NTEventLogAppender — are the only things a hand-written configuration can no longer reference.
Per-request debug tracing is off by default: it renders every request argument and every backend
response through DebugLogUtil, which costs real CPU and allocation on each RPC. Turn it on
deliberately, per environment or per incident:
log4j.logger.io.github.mmalykhin.hmsproxy=DEBUGWith it enabled, each client call gets a requestId, and the logs include:
- incoming HMS request method and arguments
- selected backend catalog
- proxied thrift method and rewritten arguments
- backend response or backend error
- final client response or client-visible error
Rendered values are bounded (10 elements per collection, 3 levels deep, ~4000 characters per
record), but the volume is still substantial on a busy proxy. Turn it back to INFO when done.
Note that INFO on the proxy package already keeps the trace stage=client-request /
backend-request write-path trace lines, so most operational debugging does not need DEBUG.
Point HiveServer2 hive.metastore.uris to this proxy instead of a single backend HMS.
For multi-catalog deployments, prefer Hive versions and clients that preserve catalog fields.
If your clients are older, use catalog<separator>db.table naming consistently, with
<separator> taken from routing.catalog-db-separator.
For Beeline/HS2 SQL workloads, a non-dot separator is usually easier to use than ..
Recommended example:
routing.catalog-db-separator=__This only changes the external legacy spelling from the canonical routing model above.
If HiveServer2 metadata writes behave differently through the proxy than directly against the backend HMS, try enabling:
federation.preserve-backend-catalog-name=trueThis only changes the returned catName/catalogName, typically to backend values such as
hive. Backend selection still follows the canonical routing model above.
If your workloads depend on Hive views or materialized views across multiple catalogs, also test with:
federation.view-text-rewrite.mode=REWRITEThat rewrites view SQL between external and internal names. viewOriginalText is preserved by
default (federation.view-text-rewrite.preserve-original-text=true); set it to false only if the
stored original SQL must be translated as well. Neither switch changes backend selection for the
RPC itself.
If you need the proxy to physically delete external-table data on Apache 3.1.3 backends after
DROP TABLE, enable:
federation.external-table-drop-purge.mode=BEST_EFFORT
catalog.catalog2.conf.hms.proxy.external-table-drop-purge.allowed-prefixes=hdfs://ns-catalog2/tmp/,hdfs://ns-catalog2/data/external/This hook currently applies only to drop_table and drop_table_with_environment_context, only
for backend runtime APACHE_3_1_3, only when the table is already EXTERNAL_TABLE with
external.table.purge=true, and only when the qualified LOCATION matches the configured
allowlist. The proxy deletes data only after the backend metastore drop succeeds; purge failures
are logged and do not roll back the metadata delete. Keep the allowlist narrow because matching
locations are deleted recursively through Hadoop FileSystem.
The recursive delete runs on a small background worker pool (threads named
hms-proxy-drop-purge-*), so drop_table returns as soon as the metadata delete succeeds and does
not hold a Thrift worker for the whole delete. A successful DROP TABLE therefore does not
guarantee that the data is already gone — watch the proxy log for the purge outcome. Pending purges
are awaited during proxy shutdown; purges still queued after that are logged and skipped.
Then a non-default catalog is used through the external database name:
USE catalog2__sales;
SHOW TABLES IN catalog2__sales;
SELECT * FROM `catalog2__sales`.orders LIMIT 10;If you keep the default separator ., older Hive SQL clients can treat catalog.db.table
ambiguously, so __ is usually the safer choice.
The proxy can also run a parallel HTTP listener that speaks the Iceberg REST
Catalog spec, backed by the same routing/federation pipeline as the Thrift HMS
front door. Status: experimental; the full write surface RESTCatalogAdapter
exposes - table writes (create, commit, drop, rename, register), view writes
(create, commit, drop, rename) and namespace DDL (create, update properties,
drop), plus multi-table transaction commit - is supported, but only when the
target namespace resolves to routing.default-catalog. Iceberg clients
(PyIceberg, Spark iceberg-rest, Trino iceberg-rest) can discover and load
Iceberg tables stored in HMS via the standard metadata_location table
parameter.
Table writes and the default-catalog-only gate landed first; view writes and
multi-table transaction commit were already reachable at that point too - the
REST dispatch path for them shares the same generic RoutingHiveCatalog/
RoutingMetaStoreClient plumbing table writes use, and WriteRouteGate
already classified all thirteen routes as writes - they were simply not yet
advertised in GET /v1/config or covered by smoke. Namespace DDL is genuinely
new: RoutingMetaStoreClient did not implement createDatabase,
alterDatabase or dropDatabase until now, so POST /v1/{prefix}/namespaces
and friends answered UnsupportedOperationException regardless of catalog.
Why writes are default-catalog only: only the default catalog's tables
are backed by a real HMS lock (see ZooKeeper storage for synthetic read
locks); every other
catalog is served by the synthetic lock shim, which grants an EXCLUSIVE
lock unconditionally with no conflict checking. A commit routed there would
believe it owns the table while silently racing - and possibly losing to - a
concurrent writer, corrupting metadata.json without ever reporting a
conflict. WriteRouteGate enforces this on the resolved catalog, not on
the request's own URL prefix: the default catalog's own prefix also exposes
every other catalog's databases as federated <catalog><separator><db>
names (see Supported endpoints below), so a create
under /v1/{default-prefix}/namespaces/apache__default/tables is refused
exactly like a direct create under /v1/apache/namespaces/default/tables -
both resolve to the apache catalog and get the same 403
(ForbiddenException) with a message naming the resolved catalog. This
applies uniformly to every one of the thirteen write routes below - table,
view and namespace DDL, and transaction commit alike. GET /v1/config and
GET /v1/{prefix}/config advertise this asymmetry directly: the default
catalog's endpoints list carries all thirteen write routes below, every
other catalog's carries only the nine read routes, so a spec-compliant client
discovers the restriction instead of learning about it from a failed request.
Enable it via:
rest-catalog.enabled=true
rest-catalog.port=9183
# Optional but recommended for production: SPNEGO. Requires security.mode=KERBEROS.
rest-catalog.kerberos.principal=HTTP/_HOST@EXAMPLE.COM
rest-catalog.kerberos.keytab=/etc/security/keytabs/spnego.service.keytab
# Optional: bound what a purge may delete (see the purge section below).
rest-catalog.purge.mode=ALLOWLIST
rest-catalog.purge.allowed-prefixes=hdfs://ns-default/warehouse/tablespace/rest-catalog.purge.allowed-prefixes is required by ALLOWLIST and rejected
with the other two modes; both contradictions fail the proxy at startup rather
than being ignored.
Requests to this listener are covered by the Prometheus metrics described in
Prometheus metrics: hms_proxy_rest_requests_total,
hms_proxy_rest_request_duration_seconds, and hms_proxy_rest_listener_info.
| Endpoint | Status |
|---|---|
GET /v1/config |
supported |
GET /v1/{prefix}/config |
supported |
GET /v1/{prefix}/namespaces |
supported |
GET /v1/{prefix}/namespaces/{ns} |
supported |
GET /v1/{prefix}/namespaces/{ns}/tables |
supported (Iceberg tables only) |
GET /v1/{prefix}/namespaces/{ns}/tables/{tbl} |
supported (Iceberg tables only) |
GET /v1/{prefix}/namespaces/{ns}/views |
supported (real listing, empty unless Iceberg views exist) |
GET /v1/{prefix}/namespaces/{ns}/views/{view} |
supported (Iceberg views only) |
HEAD /v1/{prefix}/namespaces/{ns} |
supported (204 if exists, 404 if not) |
HEAD /v1/{prefix}/namespaces/{ns}/tables/{tbl} |
supported (204 if exists, 404 if not) |
HEAD /v1/{prefix}/namespaces/{ns}/views/{view} |
supported (204 if exists, 404 if not) |
POST /v1/{prefix}/namespaces/{ns}/tables |
supported for the default catalog only (create); 403 elsewhere |
POST /v1/{prefix}/namespaces/{ns}/tables/{tbl} |
supported for the default catalog only (commit/update); 403 elsewhere |
DELETE /v1/{prefix}/namespaces/{ns}/tables/{tbl} |
supported for the default catalog only (drop, including ?purgeRequested=true); 403 elsewhere |
POST /v1/{prefix}/tables/rename |
supported for the default catalog only (rename); 403 elsewhere |
POST /v1/{prefix}/namespaces/{ns}/register |
supported for the default catalog only (register); 403 elsewhere |
POST /v1/{prefix}/namespaces/{ns}/views |
supported for the default catalog only (create); 403 elsewhere |
POST /v1/{prefix}/namespaces/{ns}/views/{view} |
supported for the default catalog only (commit/update); 403 elsewhere |
DELETE /v1/{prefix}/namespaces/{ns}/views/{view} |
supported for the default catalog only (drop); 403 elsewhere |
POST /v1/{prefix}/views/rename |
supported for the default catalog only (rename); 403 elsewhere |
POST /v1/{prefix}/namespaces |
supported for the default catalog only (create); 403 elsewhere |
POST /v1/{prefix}/namespaces/{ns}/properties |
supported for the default catalog only (update properties); 403 elsewhere |
DELETE /v1/{prefix}/namespaces/{ns} |
supported for the default catalog only (drop); 403 elsewhere |
POST /v1/{prefix}/transactions/commit |
supported for the default catalog only (multi-table commit); 403 elsewhere |
{prefix} is any catalog listed in catalogs=: every configured catalog is
exposed as its own REST prefix, /v1/<catalog>/.... GET /v1/config supports
warehouse discovery: pass ?warehouse=<catalog> and the response's
overrides.prefix names that catalog, so a client can bind itself to the
right prefix without hardcoding it (see the client examples below). Without
warehouse, /v1/config advertises routing.default-catalog, matching
phase-1 behavior; an unknown warehouse value returns HTTP 400
(BadRequestException). The response's endpoints field lists the nine
read routes above for every catalog (list/load namespace + namespace-exists,
list/load table + table-exists, list/load view + view-exists); the default
catalog's endpoints additionally carry all thirteen write routes above
(table, view and namespace DDL, transaction commit), so a modern client can
discover the write/read asymmetry between catalogs instead of learning about
it from a failed request. GET /v1/{prefix}/config
answers the same way from the proxy's own handler — overrides.prefix names
the catalog from the path segment instead of the warehouse query param —
and an unknown prefix there is still a 404.
A refused write answers 403 (ForbiddenException) with a message naming
the resolved catalog, for example: Writes are only supported in the default catalog 'hdp'; namespace 'apache__default' belongs to catalog 'apache', which is served by the synthetic lock shim and provides no writer isolation. This is enforced on the namespace the request resolves to,
not on the URL prefix it arrived under - see Why writes are default-catalog
only above.
DELETE .../tables/{tbl}?purgeRequested=true is served as a real purge, the
way Iceberg's own REST catalog serves it: the proxy drops the table in the
metastore and then deletes the table's data and metadata files itself, walking
the table's manifests to find them. That deletion runs inside the proxy's JVM
under the proxy's own credentials and finishes before the 204 is sent. The
write gate keeps it inside the default catalog: a purge whose namespace resolves
elsewhere is refused with the same 403 as any other write, before any file is
touched.
What it may delete inside that catalog is bounded by rest-catalog.purge.mode:
| Mode | Behaviour |
|---|---|
ALLOW (default) |
Deletes whatever the table's metadata and manifests point at, the way Iceberg REST catalogs do. Unchanged from earlier versions. |
ALLOWLIST |
Purges only under rest-catalog.purge.allowed-prefixes. |
REFUSE |
Answers 403 to every purge; a DELETE without purgeRequested still drops the table and leaves its files. |
A purge the policy does not permit is refused before the table is dropped,
so a refusal leaves both the table and its files intact and the client can retry
without purgeRequested.
ALLOWLIST enforces the boundary twice. Before the drop, the table's location
and its metadata.json are matched against the prefixes; either outside them
answers 403. During the deletion, every path Iceberg asks to delete — data
files, delete files, manifests, manifest lists, metadata files — is matched
again, and a path outside the prefixes is skipped and logged at WARN instead of
deleted. The second check is the one that matters for untrusted clients: in the
REST protocol the client writes the manifests, so a commit can point a snapshot
at files in another tree, and those paths only become known while walking the
manifests — after the pre-flight check has already passed.
Error responses carry the mapped HTTP status, type and message but never
a server stack trace, since this listener can be reached without
authentication. A request body that fails to parse answers 400
(BadRequestException) instead of falling through to a 500; this applies to
every route that takes a body. A HEAD response never writes a body, per RFC
9110 — including on an error status — so an exists-check against a missing
namespace, table or view returns a plain 404 with no body, not a 404 with a
JSON payload. The request dispatcher catches Throwable, not just
Exception: any Error that escapes handling (for example a dependency-
version NoSuchMethodError surfacing deep inside a write) is mapped to the
usual error response instead of unwinding past the handler and leaving the
client's connection to hang forever with no response at all.
The default catalog's prefix keeps the phase-1 federated view: its own
databases plus every other catalog's databases under their
<catalog><separator><db> names (see HiveServer2 above).
Every other prefix is a clean view: only that catalog's own databases, under
their internal names — the federated <catalog><separator><db> names never
leak into a non-default prefix. An unknown prefix still returns 404
(NoSuchCatalogException).
SPNEGO is an HTTP RFC requirement: the principal must be HTTP/<host>@REALM.
This is a separate principal from security.server-principal (which is
usually hms/<host>@REALM for the Thrift listener). Both can live in the same
keytab file or in two separate keytabs. The REST listener calls
UserGroupInformation.loginUserFromKeytabAndReturnUGI to acquire its own UGI
without overwriting the Thrift one, so they coexist in the same JVM.
PyIceberg:
from pyiceberg.catalog.rest import RestCatalog
catalog = RestCatalog("my-catalog", **{
"uri": "http://hms-proxy:9183",
})Spark:
spark.sql.catalog.my_catalog=org.apache.iceberg.spark.SparkCatalog
spark.sql.catalog.my_catalog.catalog-impl=org.apache.iceberg.rest.RESTCatalog
spark.sql.catalog.my_catalog.uri=http://hms-proxy:9183To target a non-default catalog, pass warehouse=<catalog>; the client sends
it to GET /v1/config during discovery and binds itself to that catalog's
prefix for every following request:
catalog = RestCatalog("sales-catalog", **{
"uri": "http://hms-proxy:9183",
"warehouse": "sales",
})spark.sql.catalog.sales_catalog.warehouse=sales- Iceberg REST is Iceberg-spec only by design. Non-Iceberg Hive tables
(parquet/orc/text without
metadata_location) are filtered out by HiveCatalog and stay invisible through REST. Continue to use the Thrift listener for native Hive tables. RoutingHiveCataloguses reflection on Iceberg's privateHiveCatalog.clientsfield, pinned to Iceberg1.9.2. Bumping the Iceberg version requires runningRoutingHiveCatalogTestto confirm the inject still works.- A table write opens an HDFS output stream (
metadata.json) from inside the proxy's own JVM - the first code path in the proxy to do so; reads use a different, unaffected class path. This requireshadoop-hdfsandhadoop-commonto be the same version in the dependency tree. Maven's mediation never compared them (they are different artifact IDs), soorc-core(pulled in transitively byhive-standalone-metastore) was dragging a stalehadoop-hdfs:2.2.0alongsidehadoop-common:2.6.0elsewhere in the tree, and every write failed withNoSuchMethodError: FSOutputSummer.<init>deep insideDFSOutputStream.pom.xmlnow excludes that transitivehadoop-hdfsand depends onhadoop-hdfs:2.6.0directly, to matchhadoop-common. Keep the two aligned if you ever override either version. - A purge reads the table's manifests through Avro.
iceberg-core:1.9.2is compiled againstavro:1.12.0, whilehadoop-mapreduce-client-core:2.6.5dragsavro:1.7.4in at the same depth of the tree - Maven picked 1.7.4 by declaration order and every purge died withNoClassDefFoundError: org/apache/avro/LogicalTypes.pom.xmlpinsavro:1.12.0; nothing else on this classpath calls the Avro API. Avro's snappy, xz and zstd codecs are optional dependencies and deliberately stay out of the fat jar: a manifest is always written with deflate, which Avro handles with the JDK's own Deflater.write.avro.compression-codecdoes not change that - it drives data files and delete files, which this proxy never reads, and Iceberg'sManifestWriternever applies table properties to the manifest appender. A test pins that (dropTableWithPurgeReadsManifestsOfATableAskingForSnappy), so if a future Iceberg starts honouring the property for manifests it fails there rather than in production.
The default mode. No keytab or principals required:
server.port=9083
security.mode=NONE
catalogs=warehouse
catalog.warehouse.conf.hive.metastore.uris=thrift://hms-backend:9083
routing.default-catalog=warehouseSecurity is configured in two independent parts: the front door (clients → proxy) and the backend connections (proxy → HMS).
Front door — proxy listens with SASL:
security.mode=KERBEROS
security.server-principal=hive/_HOST@REALM.COM
security.keytab=/etc/security/keytabs/hms-proxy.keytab
# Optional: principal used when connecting to backends.
# Defaults to server-principal if not set.
security.client-principal=hive/_HOST@REALM.COM
# Optional: dedicated keytab used when connecting to backend metastores.
# Defaults to security.keytab if not set.
security.client-keytab=/etc/security/keytabs/hms-proxy-client.keytabsecurity.server-principal and security.keytab are required when security.mode=KERBEROS.
The keytab must exist and be readable — the proxy will fail to start otherwise.
_HOST is replaced with the proxy machine's canonical hostname before Kerberos login. If your DNS
canonical name differs from the keytab/KDC principal name, use the explicit FQDN instead.
When Kerberos is enabled on the front door, delegation-token methods
(get_delegation_token, renew_delegation_token, cancel_delegation_token)
are served locally by the proxy's own token manager instead of being forwarded
to backend metastores.
That also means proxy-user authorization for service principals such as HiveServer2 must be
configured on the proxy front door itself, not only on backend HMS instances. If HiveServer2
connects as hive/_HOST@REALM.COM and requests get_delegation_token("alice", ...), the proxy
must see matching hadoop.proxyuser.hive.* rules through core-site.xml or
security.front-door-conf.* overrides.
The proxy reads delegation-token store settings from the usual Hive configuration
resources visible to the proxy process through HiveConf, typically hive-site.xml.
You can also set them directly in hms-proxy.properties via
security.front-door-conf.<hive-conf-key>=..., which is often simpler when the
launcher script does not put hive-site.xml on the proxy classpath.
If no persistent token store is configured, Hive falls back to
org.apache.hadoop.hive.metastore.security.MemoryTokenStore, and then a proxy
restart invalidates existing HiveServer2 delegation tokens.
hadoop.proxyuser.<service>.* is not related to ZooKeeper transport. It is only needed when a
service principal such as HiveServer2 asks the proxy to issue an end-user delegation token via
get_delegation_token("alice", ...).
Example:
# Allow HiveServer2's Kerberos principal to request end-user delegation tokens from the proxy.
security.front-door-conf.hadoop.proxyuser.hive.hosts=hs2-1.example.com,hs2-2.example.com
security.front-door-conf.hadoop.proxyuser.hive.groups=*Example with ZooKeeperTokenStore configured directly in the proxy config:
security.front-door-conf.hive.cluster.delegation.token.store.class=org.apache.hadoop.hive.metastore.security.ZooKeeperTokenStore
security.front-door-conf.hive.cluster.delegation.token.store.zookeeper.connectString=zk1:2181,zk2:2181,zk3:2181
security.front-door-conf.hive.cluster.delegation.token.store.zookeeper.znode=/hms-proxy-delegation-tokens
# Optional: ACL applied to newly created ZooKeeper nodes for the token store.
# security.front-door-conf.hive.cluster.delegation.token.store.zookeeper.acl=sasl:hive:cdrwa
# Optional: token max lifetime in milliseconds.
# security.front-door-conf.hive.cluster.delegation.token.max-lifetime=604800000When ZooKeeperTokenStore is enabled and security.mode=KERBEROS, the proxy now also
auto-populates hive.metastore.kerberos.principal and hive.metastore.kerberos.keytab.file
for the front-door HiveConf from security.server-principal and security.keytab. That lets
Hive's built-in ZooKeeperTokenStore client authenticate to ZooKeeper over SASL/Kerberos
without requiring a separate JAAS file in the common case. If you need different credentials for
ZooKeeper, set those hive.metastore.kerberos.* keys explicitly through security.front-door-conf.*.
In other words, the default ZooKeeper client credentials are:
- principal:
security.server-principal - keytab:
security.keytab
These are the front-door service credentials. They are not taken from
security.client-principal or security.client-keytab, which are used for outbound backend HMS
connections instead.
At startup the proxy also pre-configures Hive's HiveZooKeeperClient JAAS entry for the
front-door token store from those same credentials before starting the local delegation-token
manager. In the usual setup that means you do not need a separate -Djava.security.auth.login.config
file just for ZooKeeper SASL.
If ZooKeeper must use different credentials from the proxy front door, override them explicitly:
security.front-door-conf.hive.metastore.kerberos.principal=hive-zk/_HOST@REALM.COM
security.front-door-conf.hive.metastore.kerberos.keytab.file=/etc/security/keytabs/hms-proxy-zk.keytabOn startup the proxy logs which HiveConf resources it found and which front-door
overrides were applied. If you see Found configuration file null from Hive or the
proxy warns that it is using MemoryTokenStore, the process did not see a persistent
delegation-token store config. If get_delegation_token fails with
User: hive/... is not allowed to impersonate alice, the proxy is missing front-door
hadoop.proxyuser.<service>.hosts/groups rules for that Kerberos caller.
The proxy also has a narrow synthetic lock shim for non-default catalogs. It covers
non-ACID SHARED_READ SELECT locks, eligible non-transactional NO_TXN DDL locks
such as CREATE TABLE and partition rename/drop flows that Hive still runs through the
txn/lock APIs, and non-transactional write locks (INSERT, UPDATE, DELETE) — the locks
Hive takes for an INSERT into a non-ACID table of a non-default catalog. The proxy serves
those locks locally when backend txn ids do not line up across catalogs. This does not turn
the proxy into a distributed ACID coordinator.
Write locks are not restricted by lock type: with Hive's default
hive.txn.strict.locking.mode=true an INSERT into a non-ACID table takes an EXCLUSIVE
lock, SHARED_WRITE only when strict locking is off, and INSERT OVERWRITE is always
EXCLUSIVE. A component that declares isTransactional=true is never synthesized: it belongs
to an ACID table, whose write ids only the default catalog's TxnHandler can allocate, so the
request is left to the backend and fails there. ACID tables cannot exist in a non-default
catalog anyway — the proxy rejects create_table with transactional=true outside the default
catalog and refuses allocate_table_write_ids / get_valid_write_ids for non-default catalogs.
The shim grants locks without checking whether they conflict. It answers ACQUIRED
immediately and never inspects other live locks, so concurrent INSERT, INSERT OVERWRITE and
DDL statements against the same table of a non-default catalog are not serialized against each
other, and neither are clients that reach that metastore without passing through the proxy. A
non-default catalog therefore gives no ACID or writer-isolation guarantees; if a workload needs
mutual exclusion, keep it in routing.default-catalog or serialize the writers yourself. The
catalog access mode is still enforced: a write lock for a READ_ONLY catalog, or for a database
outside catalog.<name>.write-db-whitelist, is rejected with a MetaException instead of being
synthesized.
Eligibility is decided over every LockComponent of the request, not just the first one. A
lock call is acquired, routed, and acknowledged as a single unit and yields one lock id, so it
cannot be forwarded to more than one backend. A query that reads across catalogs - or merely
across two databases of one catalog - nevertheless arrives as exactly that: one request whose
components resolve to different namespaces. The proxy splits it. Components of one catalog are
routed to that catalog's metastore, each rewritten to its own backend database, and components of
the other catalogs are dropped from the request that reaches the backend.
The default catalog is chosen as the routing target whenever it is present, because it owns the
TxnHandler and its locks are the real ones. Dropping the rest costs nothing that was ever held:
a non-default catalog is served by the shim described above, which records a lock and reports it
as acquired without ever testing for a conflict. What a dropped component loses is an entry in
that ledger, not a guarantee. Access-mode enforcement does not travel with the routing decision -
a write into a READ_ONLY catalog is still refused whether or not its component survived the
split. Each split is logged and counted by hms_proxy_lock_request_split_total.
Hive's INSERT ... VALUES placeholder (_dummy_database/_dummy_table, the
SemanticAnalyzer.DUMMY_DATABASE constant) is the one exception: it exists in no metastore and
belongs to no catalog, so its components neither select a catalog, nor count as a second
namespace, nor get rewritten into a real backend database. A request made only of placeholder
components goes to the default catalog that owns the TxnHandler.
Hive's _dummy_database._dummy_table pseudo source is the one exception: INSERT ... VALUES
and FROM-less queries lock it next to the real target table, so such a request always names two
databases. That pseudo table exists in no metastore and there is nothing to lock on it, so the
proxy skips it when it decides both the namespace and the shim eligibility of a lock request. A
request that names only the pseudo source still follows the default-catalog pin.
synthetic-read-lock.store.mode must be set explicitly — there is no default. Use IN_MEMORY
for single-instance deployments (non-default catalog synthetic locks are lost on proxy restart or
load-balancer failover, and the proxy logs a WARN on startup to surface that), or ZOOKEEPER
for HA / load-balanced deployments so check_lock, unlock, heartbeat, commit_txn, and
abort_txn can continue through a different proxy instance after the first one dies.
mode=IN_MEMORY together with any configured synthetic-read-lock.store.zookeeper.* property is a
startup error: the ZooKeeper settings would be ignored and locks would silently stay in memory. The
opposite direction is preserved — ZooKeeper properties without an explicit mode select
ZOOKEEPER.
Example:
synthetic-read-lock.store.mode=ZOOKEEPER
synthetic-read-lock.store.zookeeper.connect-string=zk1:2181,zk2:2181,zk3:2181
synthetic-read-lock.store.zookeeper.znode=/hms-proxy-synthetic-read-locks
# synthetic-read-lock.store.zookeeper.connection-timeout-ms=15000
# synthetic-read-lock.store.zookeeper.session-timeout-ms=60000
# synthetic-read-lock.store.zookeeper.base-sleep-ms=1000
# synthetic-read-lock.store.zookeeper.max-retries=3When security.mode=KERBEROS, the synthetic read-lock store uses the same
security.server-principal and security.keytab credentials for ZooKeeper SASL/Kerberos
as the front door by default, similar to the delegation-token store setup above.
The store keeps two subtrees under the configured znode:
/hms-proxy-synthetic-read-locks/locks/lock-<sequence> serialized lock state, keyed by lock id
/hms-proxy-synthetic-read-locks/txns/<txnId>/<lockId> empty pointer node, keyed by transaction
check_lock, unlock, and heartbeat only carry a lock id, so lock state has to stay keyed by
lock id; the txns subtree is the secondary index that lets commit_txn and abort_txn release a
transaction's locks with a single getChildren instead of scanning every live lock. A commit for a
transaction that never took a synthetic lock costs one getChildren that returns NoNode.
Expired locks are collected by a background sweep every 30s (never on a client request thread), and
that sweep also removes index entries whose lock node is gone and empty per-transaction parents.
Locks are always written before their index entry, so a half-finished create can never hide a lock
from commit_txn.
Rolling upgrade: locks written by a proxy older than this layout have no index entry, so a newer
instance will not release them on commit_txn/abort_txn — they expire through the sweep within
metastore.txn.timeout instead. Synthetic locks do not block anything on the backend, so the only
cost is a znode that lives slightly longer. No manual migration or downtime is required, and mixed
old/new instances can run against the same znode.
The lock timeout is derived from metastore.txn.timeout of the default catalog and accepts Hive's
canonical suffixed form (600s, 10m); a bare number is still read as seconds. An unparsable value
falls back to 300s and is logged as a WARN at startup.
If you want backend HMS calls to execute as the authenticated front-door Kerberos user instead of the proxy service principal, enable:
security.impersonation-enabled=trueWhen enabled, the proxy derives the caller identity from the inbound Kerberos/SASL session and
keeps a separate cached backend HMS client per user and per catalog, issuing set_ugi() once when
that cached client is opened. This avoids leaking one user's identity into another user's requests
while also avoiding a full backend reconnect on every RPC.
You can also override this per backend:
# Global default for catalogs that do not specify their own setting.
security.impersonation-enabled=false
catalog.catalog1.impersonation-enabled=true
catalog.catalog2.impersonation-enabled=falseThis lets you enable caller impersonation only for selected backends while leaving the others on the proxy service principal. The global key acts purely as that default: at runtime impersonation is driven by the per-catalog flag, which inherits the global value when it is not set explicitly.
This mode requires security.mode=KERBEROS on the proxy listener. If a legacy client explicitly
calls set_ugi, the proxy will ignore the requested username and use the authenticated Kerberos
caller instead.
If internal HiveServer2 traffic works as user hive, but your personal Kerberos user fails even
with admin SQL privileges, that usually means backend impersonation is working as designed:
- service traffic from
hiveis allowed to continue on the proxy service principal - real end users are sent to the backend through
set_ugi()and therefore need backend Hadoop proxy-user permission forsecurity.client-principal - SQL admin rights in Hive do not replace Hadoop proxy-user authorization
In that case either allow proxy-user impersonation on the backend HMS for the proxy outbound
principal, or disable impersonation for that specific backend with
catalog.<name>.impersonation-enabled=false.
Backend connections — you can set shared defaults for all backend thrift clients via
backend.conf.*, then override per catalog via catalog.<name>.conf.*:
backend.conf.hive.metastore.client.socket.timeout=45s
backend.conf.hive.metastore.execute.setugi=true
backend.conf.hive.metastore.uris=thrift://shared-hms.internal:9083Per-catalog values win over shared backend defaults:
catalog.catalog1.conf.hive.metastore.uris=thrift://hms-a.internal:9083
catalog.catalog1.conf.hive.metastore.sasl.enabled=true
catalog.catalog1.conf.hive.metastore.kerberos.principal=hive/_HOST@REALM.COMPer catalog you can also restrict write traffic:
catalog.catalog1.access-mode=READ_ONLY
catalog.catalog2.access-mode=READ_WRITE_DB_WHITELIST
catalog.catalog2.write-db-whitelist=sales,analyticsPer catalog you can also define a latency budget for the latency-aware routing layer:
catalog.catalog1.latency-budget-ms=1500
catalog.catalog2.latency-budget-ms=5000And you can independently restrict which metadata is visible from each backend:
catalog.catalog1.expose-mode=DENY_BY_DEFAULT
catalog.catalog1.expose-db-patterns=sales
catalog.catalog1.expose-table-patterns.sales=orders_.*,eventsSupported modes:
READ_WRITE: default behaviorREAD_ONLY: only read RPCs are allowed for that catalogREAD_WRITE_DB_WHITELIST: writes are allowed only when the resolved backend database is listed incatalog.<name>.write-db-whitelist
access-mode and write-db-whitelist must be configured together. A whitelist under any other
access mode, or READ_WRITE_DB_WHITELIST without a whitelist, is a startup error rather than a
catalog that silently allows every write.
Exposure modes:
ALLOW_ALL: default metadata exposure behavior forcatalog.<name>.expose-modeDENY_BY_DEFAULT: metadata is hidden unless matched bycatalog.<name>.expose-db-patternsorcatalog.<name>.expose-table-patterns.<dbRegex>
This means you can add arbitrary HiveConf keys used by HiveMetaStoreClient startup, not only
the metastore URI and Kerberos settings.
Latency-aware routing — optional backend-state polling, adaptive timeouts, circuit breaking,
safe hedged fanout reads, and degraded omission are configured with routing.* properties:
routing.backend-state-polling.enabled=true
routing.backend-state-polling.interval-ms=10000
routing.backend-state-polling.probe-timeout-ms=5000
routing.backend-state-polling.max-parallelism=8
routing.adaptive-timeout.enabled=true
routing.adaptive-timeout.initial-ms=5000
routing.adaptive-timeout.min-ms=1000
routing.adaptive-timeout.max-ms=60000
routing.adaptive-timeout.multiplier=4.0
routing.adaptive-timeout.alpha=0.2
routing.adaptive-timeout.reconnect-cooldown-ms=30000
routing.circuit-breaker.enabled=true
routing.circuit-breaker.failure-threshold=3
routing.circuit-breaker.open-state-ms=30000
routing.hedged-read.enabled=true
routing.hedged-read.max-parallelism=8
routing.degraded-routing-policy=SAFE_FANOUT_READSThe adaptive timeout starts from routing.adaptive-timeout.initial-ms, then follows backend
latency EWMA within the configured min/max bounds. Each timeout change reconnects the shared
backend client and evicts cached impersonation clients (Kerberos re-login), so the proxy
applies hysteresis (max(2s, 25 % of the current value)) and a cooldown
(routing.adaptive-timeout.reconnect-cooldown-ms, default 30 s) before reconnecting again. The
counters hms_proxy_adaptive_timeout_reconnect_total and
hms_proxy_adaptive_timeout_reconnect_skipped_total{reason="hysteresis"|"cooldown"} expose how
often reconnects fire vs. are suppressed. A reconnect that cannot quiesce the shared pool in time
rolls the client socket timeout back to the value the live sessions use and still starts the
cooldown, so an overloaded backend is not re-quiesced on every following request. Transport
failures and latency-budget breaches
count toward the circuit breaker. Once a backend crosses routing.circuit-breaker.failure-threshold,
the proxy fails fast for that backend until the open window expires, then lets one half-open retry
decide whether to close or reopen the circuit.
Iceberg pointer guard — a HiveServer2 INSERT into an Iceberg table opens with an
alter_table_with_environment_context carrying the Table the query snapshotted at compile time,
and a metastore applies those parameters wholesale, so every Iceberg key the record holds and the
request omits is erased - metadata_location first of all. The proxy is the only place the Hive
and Iceberg write paths meet, so it merges such an alter over the record the metastore currently
holds instead of forwarding it as sent; an honest Iceberg commit, recognised by a
previous_metadata_location equal to the current pointer, is passed through untouched. On Hive 4
backends the repaired alter also carries expected_parameter_key/expected_parameter_value, which
makes the metastore apply it only while the pointer is still the one that was read.
routing.iceberg-pointer-guard.enabled=true
routing.iceberg-pointer-guard.table-cache-ttl-ms=30000
routing.iceberg-pointer-guard.table-cache-max-entries=10000
routing.iceberg-pointer-guard.lock-enabled=true
routing.iceberg-pointer-guard.lock-acquire-timeout-ms=10000Reading the record and writing the alter are two calls, so a repair holds the table lock Iceberg
itself takes (EXCLUSIVE, table level - the request shape of
org.apache.iceberg.hive.MetastoreLock) across the re-read, the merge and the backend's
alter_table. The lock is requested only when a repair is needed: an honest Iceberg commit
sends its alter_table from inside that very lock, so asking for it before deciding would wait on
the caller of the call being served. A lock that is not granted within lock-acquire-timeout-ms
never refuses the write - the alter goes through repaired but unprotected, counted as
repair_lock_timeout - because failing an ordinary Hive write whenever the metastore's lock table
hiccups is worse than the race being closed. 0 means one attempt and no waiting. Non-default
catalogs are never locked: their writers are served by the synthetic lock shim, which grants locks
without checking conflicts, so nothing contends for the object the lock would hold.
Whether a table is an Iceberg table is decided from the metastore's record, not from the request -
the shape HiveServer2 sends carries no Iceberg key at all - which costs one get_table per
alter_table. table-cache-ttl-ms bounds that: a name the metastore answered is not an Iceberg
table is remembered for that long, so ordinary Hive tables (where the alter_table volume is) pay
one read and then nothing. Iceberg tables are never cached, because their current pointer has to be
read fresh; 0 disables the cache entirely. A table that becomes an Iceberg table outside the
proxy is unprotected for at most one TTL - a create_table or an alter_table that carries a
pointer drops the remembered answer immediately. hms_proxy_iceberg_pointer_guard_events_total
exposes both the repairs and the reads the cache saved.
When hive.metastore.sasl.enabled=true is set for any catalog, the proxy opens outbound HMS
connections using security.client-principal and security.client-keytab. If those are omitted,
it falls back to security.server-principal and security.keytab.
Front and backend security are independent: you can run the front door without Kerberos and
still connect to Kerberos-protected backends, or vice versa. In that case, set
security.mode=NONE and provide security.client-principal plus security.client-keytab.
Full example:
server.port=9083
server.bind-host=0.0.0.0
security.mode=KERBEROS
security.server-principal=hive/_HOST@REALM.COM
security.keytab=/etc/security/keytabs/hms-proxy.keytab
security.client-principal=hive/_HOST@REALM.COM
security.client-keytab=/etc/security/keytabs/hms-proxy-client.keytab
catalogs=catalog1,catalog2
routing.default-catalog=catalog1
routing.backend-state-polling.enabled=true
routing.adaptive-timeout.enabled=true
routing.circuit-breaker.enabled=true
routing.hedged-read.enabled=true
routing.degraded-routing-policy=SAFE_FANOUT_READS
catalog.catalog1.conf.hive.metastore.uris=thrift://hms-a.internal:9083
catalog.catalog1.conf.hive.metastore.sasl.enabled=true
catalog.catalog1.conf.hive.metastore.kerberos.principal=hive/_HOST@REALM.COM
catalog.catalog1.latency-budget-ms=1500
catalog.catalog2.conf.hive.metastore.uris=thrift://hms-b.internal:9083
catalog.catalog2.conf.hive.metastore.sasl.enabled=true
catalog.catalog2.conf.hive.metastore.kerberos.principal=hive/_HOST@REALM.COM
catalog.catalog2.latency-budget-ms=5000For the scenarios in SMOKE.md, the repo now includes a runnable direct HMS API smoke client:
io.github.mmalykhin.hmsproxy.tools.HmsMetastoreSmokeCli txnio.github.mmalykhin.hmsproxy.tools.HmsMetastoreSmokeCli lockio.github.mmalykhin.hmsproxy.tools.HmsMetastoreSmokeCli notification
When a txn or lock run fails after open_txns, the client issues a best-effort abort_txn
before propagating the error, --close-txn none included: a failed smoke run must not leave a
transaction holding back the ValidTxnList watermark until the heartbeat timeout expires.
The current smoke coverage is summarized in the SMOKE.md test matrix by:
- client version
- front-door profile
- backend profile
- auth mode
- method families
- expected result
Build the jar first:
mvn -DskipTests packageFor Java 17+ with Kerberos and Hadoop 2.x libraries, use the same JVM flags as the proxy:
java \
--add-opens=java.security.jgss/sun.security.krb5=ALL-UNNAMED \
--add-exports=java.security.jgss/sun.security.krb5=ALL-UNNAMED \
-cp target/hms-proxy-$(mvn -q -DforceStdout help:evaluate -Dexpression=project.version)-fat.jar \
io.github.mmalykhin.hmsproxy.tools.HmsMetastoreSmokeCli txn \
--uri thrift://proxy-host:9083 \
--auth kerberos \
--server-principal hive/proxy-host.example.com@REALM.COM \
--client-principal alice@REALM.COM \
--keytab /etc/security/keytabs/alice.keytab \
--krb5-conf /etc/krb5.conf \
--db hdp__default \
--table smoke_txn_tblThat mode performs:
open_txnsallocate_table_write_idslockcheck_lockget_valid_write_idscommit_txn
The lock mode is for direct lock lifecycle smoke, especially the synthetic shim on
non-default catalogs. It opens one txn, acquires the requested lock, runs check_lock,
optionally sends heartbeat, optionally calls unlock, and then aborts or commits the txn.
Example CREATE TABLE-style DB lock on a non-default catalog:
java \
--add-opens=java.security.jgss/sun.security.krb5=ALL-UNNAMED \
--add-exports=java.security.jgss/sun.security.krb5=ALL-UNNAMED \
-cp target/hms-proxy-$(mvn -q -DforceStdout help:evaluate -Dexpression=project.version)-fat.jar \
io.github.mmalykhin.hmsproxy.tools.HmsMetastoreSmokeCli lock \
--uri thrift://proxy-host:9083 \
--db apache__default \
--lock-type SHARED_READ \
--lock-level DB \
--operation-type NO_TXN \
--transactional falseExample partition rename/drop style lock on a non-default catalog:
java \
--add-opens=java.security.jgss/sun.security.krb5=ALL-UNNAMED \
--add-exports=java.security.jgss/sun.security.krb5=ALL-UNNAMED \
-cp target/hms-proxy-$(mvn -q -DforceStdout help:evaluate -Dexpression=project.version)-fat.jar \
io.github.mmalykhin.hmsproxy.tools.HmsMetastoreSmokeCli lock \
--uri thrift://proxy-host:9083 \
--db apache__default \
--table smoke_managed_tbl \
--partition p=2026-04-01 \
--lock-type EXCLUSIVE \
--lock-level PARTITION \
--operation-type NO_TXN \
--transactional falseThe notification mode is for Hortonworks-only add_write_notification_log, so it also needs
the HDP standalone metastore jar:
java \
--add-opens=java.security.jgss/sun.security.krb5=ALL-UNNAMED \
--add-exports=java.security.jgss/sun.security.krb5=ALL-UNNAMED \
-cp target/hms-proxy-$(mvn -q -DforceStdout help:evaluate -Dexpression=project.version)-fat.jar \
io.github.mmalykhin.hmsproxy.tools.HmsMetastoreSmokeCli notification \
--uri thrift://proxy-host:9083 \
--auth kerberos \
--server-principal hive/proxy-host.example.com@REALM.COM \
--client-principal alice@REALM.COM \
--keytab /etc/security/keytabs/alice.keytab \
--krb5-conf /etc/krb5.conf \
--db hdp__default \
--table smoke_txn_tbl \
--txn-id 1001 \
--write-id 2001 \
--files-added hdfs:///warehouse/tablespace/managed/hive/smoke_txn_tbl/delta_1001_1001_0000/bucket_00000 \
--hdp-standalone-metastore-jar /opt/hms-proxy/hive-metastore/hive-standalone-metastore-3.1.0.3.1.0.0-78.jarUseful notes:
--server-principalmust be the proxy front-door principal, not the backend HMS principal--client-principaland--keytabare the Kerberos credentials used by this smoke client- extra HiveConf overrides can be passed via repeated
--conf key=value lockis the quickest way to reproduce non-default catalogNO_TXNshim cases such asCREATE TABLEand partition rename/drop without going through Beelinenotificationshould succeed for a Hortonworks-routed catalog and fail withrequires a Hortonworks backend runtimefor an Apache-routed catalog
For repeated checks on a real proxy installation, use separate runners for simple and kerberos
front doors:
They wrap the same smoke client and run a fail-fast scenario of:
- optional Beeline / HiveServer2 SQL smoke from SMOKE.md
- direct txn/ACID smoke
- non-default catalog DB lock smoke
- optional partition lock smoke
- optional Hortonworks notification smoke
Start from the matching example config:
cp scripts/hms-real-installation-smoke.simple.env.example scripts/hms-real-installation-smoke.simple.env
cp scripts/hms-real-installation-smoke.kerberos.env.example scripts/hms-real-installation-smoke.kerberos.envThen edit the relevant HMS_SMOKE_* values and run:
scripts/run-real-installation-smoke-simple.sh --scenario all
scripts/run-real-installation-smoke-kerberos.sh --scenario allOr run a narrower slice:
scripts/run-real-installation-smoke-simple.sh --scenario sql
scripts/run-real-installation-smoke-simple.sh --scenario locks
scripts/run-real-installation-smoke-kerberos.sh --scenario notificationNotes:
- by default the runner picks the newest
target/hms-proxy-*-fat.jar - you can override the jar path with
HMS_SMOKE_FAT_JAR - if
HMS_SMOKE_BEELINE_JDBC_URLis configured,allalso runs the Beeline / HiveServer2 SQL smoke fromSMOKE.md - SQL smoke uses
HMS_SMOKE_HDP_READ_TABLE/HMS_SMOKE_APACHE_READ_TABLE, validates view rewrite and permanent UDFs by default, and can optionally run transactional SQL and materialized-view checks - set
HMS_SMOKE_SQL_RUN_VIEW_REWRITE=falseif the proxy is intentionally started withoutfederation.view-text-rewrite.mode=REWRITE; setHMS_SMOKE_SQL_RUN_UDF=falseor overrideHMS_SMOKE_SQL_UDF_CLASStogether withHMS_SMOKE_SQL_UDF_EXPECTED_RESULTwhen the HS2 classpath differs - if
HMS_SMOKE_TXN_SECONDARY_DBandHMS_SMOKE_TXN_SECONDARY_TABLEare set, the runner executes a second direct txn smoke target - if
HMS_SMOKE_NOTIFICATION_*is not configured, the notification step is skipped inall - if
HMS_SMOKE_NOTIFICATION_NEGATIVE_DBandHMS_SMOKE_NOTIFICATION_NEGATIVE_TABLEare set, the runner also executes the negative Apache-backend notification check fromSMOKE.md - the simple runner auto-loads
scripts/hms-real-installation-smoke.simple.envwhen it exists - the Kerberos runner auto-loads
scripts/hms-real-installation-smoke.kerberos.envwhen it exists