Catalog options
ClarityQ supports two Databricks catalog types:- Unity Catalog: the modern, recommended governance layer for Databricks. Supports both serverless and classic SQL warehouses.
- Legacy Hive Metastore: the original
hive_metastorecatalog. Requires a classic (cluster-based) SQL warehouse, as serverless SQL warehouses do not support the legacy Hive Metastore.
Setup
Step 1: Get Databricks connection details
To connect ClarityQ to your Databricks workspace using a service principal, you’ll need the following information:- Server hostname
- HTTP path
- Service principal Application ID
- Service principal OAuth secret
Step 2: Find your Databricks connection details
Server Hostname
- Log in to your Databricks workspace
- In the sidebar, click SQL Warehouses (or Compute → SQL Warehouses)
- Click on your SQL warehouse name
- Go to the Connection Details tab
- Copy the Server hostname (e.g.,
your-workspace.cloud.databricks.com)
HTTP Path
- In the same Connection Details tab of your SQL warehouse
- Copy the HTTP path (e.g.,
/sql/1.0/warehouses/abc123def456)
If you plan to use the Legacy Hive Metastore, make sure your HTTP path points to a classic (cluster-based) SQL warehouse, not a serverless one. Serverless SQL warehouses do not support the legacy Hive Metastore.
Step 3: Create service principal credentials
Create a dedicated service principal for ClarityQ (recommended for production environments):Create service principal for ClarityQ
- In your Databricks workspace, go to Settings → Identity and access
- Click the Service principals tab
- Click Add service principal
- Enter name:
clarityq-service-principal - Click Add
- Click on the newly created service principal
- Copy the Application ID - you’ll need this for the ClarityQ connection, where it will be entered in the Client ID input field (UUID format like
065e73cc-a713-4ae0-8212-ba7f98e32bee) - Go to the Secrets tab
- Click Generate secret
- Set expiration date (recommended for security)
- Click Generate and copy the secret immediately
- In the same service principal details page, go to Permissions tab
- Click Grant access or Add permissions
- Search for and select Service Principal: Use
- Save the changes
- This allows the service principal to execute queries (required for ClarityQ)
Grant SQL Warehouse access
You need to grant the service principal permission to use the SQL Warehouse:- In the sidebar, click SQL Warehouses (or Compute → SQL Warehouses)
- Click on your SQL warehouse name
- Go to the Permissions tab
- Search for and select
clarityq-service-principal - In the permission dropdown, select Can use
- Click Save
Grant data access permissions to ClarityQ
Note: These steps require workspace admin or account admin permissions. If you don’t have these permissions, ask your Databricks administrator to perform these steps. Choose the tab that matches your catalog type:- Unity Catalog
- Legacy Hive Metastore
Grant SELECT permissions to your ClarityQ service principal for the data you want to analyze:Step 1: Navigate to your catalog
- Go to Catalog in your Databricks workspace
- Choose the catalog containing data you want ClarityQ to access
- Select the Permissions tab
- Click the Grant button
- Search for and select your
clarityq-service-principal - Check the SELECT privilege
- Click Grant to save
- Whole catalog: Grants ClarityQ access to all schemas and tables in the catalog
- Specific schema: Navigate to the schema level and repeat the above steps to grant access to specific schemas only
- Specific tables: Navigate to individual tables and grant access to only the tables you want ClarityQ to analyze.
information_schema so ClarityQ can discover available tables and columns.Step 3b (optional): Register ClarityQ as an OAuth app for per-user permissions
Skip this step unless you want each person’s queries to run under their own Databricks identity (see Run queries as the asking user below). It requires a Databricks account admin.1
Open App connections
In the Databricks account console (not the workspace settings), go to Settings → App connections and click Add connection. Only account admins see this page.
2
Name the application and set the redirect URL
- Application Name:
ClarityQ - Redirect URLs, one per line:
https://app.clarityq.ai/api/v1/integrations/data-warehouse/oauth/callback
3
Add the access scopes
Use Add scope to add
sql and offline_access. sql lets ClarityQ run statements on your SQL warehouses; offline_access issues the refresh token that keeps a person connected between sessions. The openid, email and profile scopes may already be listed; they are harmless. all-apis is not needed.4
Client secret and refresh tokens
- Keep Generate a client secret checked. ClarityQ authenticates to Databricks as a confidential client.
- Enable single-use refresh tokens is optional. With it on, every refresh issues a new refresh token, so people who use ClarityQ regularly are never asked to reconnect; only someone idle for longer than the refresh token lifetime is. ClarityQ handles the rotation.
5
Set token lifetimes
- Access token TTL: the default 60 minutes is fine (allowed range 5 to 1,440).
- Refresh token TTL: the default 10,080 minutes is 7 days. When it lapses without use, the person reconnects and their scheduled tasks stop until they do. Raise it, up to 129,600 minutes (90 days), if that is too short for your team.
6
Add and copy the credentials
Click Add, then copy the Client ID and the Client secret. The secret is shown once; you will paste both into the ClarityQ connection form as OAuth App Client ID and OAuth App Client Secret.
SELECT on the data they may see. Their Databricks email must match their ClarityQ login email.
Step 4: Configure connection in ClarityQ
In the ClarityQ interface, fill out the connection form with the following fields:Required Fields
- Connection Name: Choose a name for this connection (e.g., “Production Databricks”)
- Server Hostname: Your Databricks server hostname (e.g.,
your-workspace.cloud.databricks.com) - HTTP Path: Your SQL warehouse HTTP path (e.g.,
/sql/1.0/warehouses/abc123def456) - Client ID: Your service principal Application ID (UUID format, e.g.,
065e73cc-a713-4ae0-8212-ba7f98e32bee) - Client Secret: The service principal OAuth secret generated in Step 3
- Catalog type: Select Unity Catalog (default) or Legacy Hive Metastore
- Choose Legacy Hive Metastore only if your data resides in the original
hive_metastorecatalog and you are using a classic (cluster-based) SQL warehouse
- Choose Legacy Hive Metastore only if your data resides in the original
Run queries as the asking user (per-user permissions)
Leave this off and every query runs as the service principal above: Databricks sees one identity, and ClarityQ’s own permissions decide what people can reach. Turn it on and ClarityQ runs each person’s queries under their own Databricks identity. Unity Catalog then applies that user’s grants, row filters and column masks, and your Databricks audit log records the real asker rather than the service principal. Two people asking the same question can legitimately get different rows. Databricks has no way for one identity to act as another, so this works through OAuth: each person signs in to Databricks once from ClarityQ and allows it to run SQL on their behalf. That is why the toggle asks for three more fields, taken from Step 3b:- OAuth App Client ID
- OAuth App Client Secret
- Refresh token TTL (minutes) — copy the value from the OAuth app. Leave the default 10080 (7 days) if you did not change it there. ClarityQ uses it to warn people two days before their sign-in lapses.
- Each user connects their Databricks account once, from the Connect your Databricks account card Ask Anything shows on a new chat, or under Settings → Apps → Data Warehouse (see Connecting Your Data Warehouse Account). Until they do, nothing runs for them on this connection. ClarityQ never falls back to the service principal for a person’s query.
- Cached results are kept separate per user, so nobody is served rows that were filtered for someone else.
- Shared dashboards resolve per viewer: the author sees their rows, a colleague with wider access sees theirs.
- Scheduled tasks run under their owner’s connected account. If that account’s sign-in expires, the task fails until the owner reconnects.
- Schema discovery, sample values and other background work have no asking user, so they continue to run as the service principal. Its grants decide what the Context Layer can see.
Step 5: Test the connection
Verify the connection in ClarityQ to ensure:- Databricks workspace accessibility
- SQL warehouse availability
- Token authentication
- Catalog and schema access
- Query execution capabilities
Connection parameters reference
Required parameters
Optional parameters
Troubleshooting
Common connection issues
Authentication failed- Verify your service principal OAuth secret is correct and not expired
- Check if the service principal has necessary permissions through Admin Settings
- Ensure the user/service principal exists and is active
- Most common cause: Service principal not added to workspace yet
- Ask your Databricks administrator to add the service principal to the workspace
- Missing permissions: Ask your admin to use the role-based approach described above
- Insufficient permissions: Ensure ClarityQ has Service Principal Use role + read access to desired data
- SQL Warehouse access denied: Ensure the service principal has Can use permission on the SQL Warehouse (see Step 3)
- A connect prompt in a chat or on a dashboard: the connection runs queries as the asking user and this person has not connected yet, or their sign-in expired. They connect from the Connect Databricks button in the chat, or under Settings → Apps → Data Warehouse.
- “Please sign in to Databricks as …”: the Databricks account used on the consent screen is not the ClarityQ user’s own. Sign out of Databricks and connect again with the matching email.
- User can connect but gets
Access Deniedor no rows: with per-user permissions on, it is that user’s grants that apply, not the service principal’s. Check theirSELECTgrants, Can use on the warehouse, and your row filters. - Row filters not applied: Unity Catalog row filters and column masks apply to tables, not views, and only on Unity Catalog tables on a SQL warehouse. Nothing in
hive_metastoreis filtered. - Scheduled task suddenly fails: the owner’s refresh token expired (7 days by default on the OAuth app). The owner reconnects, or the account admin raises the refresh token lifetime.
- “Databricks did not return a refresh token”: the OAuth app is missing the
offline_accessscope.
- Tables not visible: Ensure
READ_METADATAhas been granted on the schema. Without it, ClarityQ can connect but cannot discover tables. - Serverless warehouse errors: The legacy Hive Metastore is not supported on serverless SQL warehouses. Switch to a classic (cluster-based) SQL warehouse.
- Permission denied on GRANT statements: Hive Metastore grants require workspace admin or the
GRANTprivilege. Ask your administrator to run the SQL grant statements.