Guidewire Cloud Data Access Private Preview
Guidewire Cloud Data Access (CDA) is a streaming export service within the Guidewire Data Platform that captures changes from InsuranceSuite applications (PolicyCenter, BillingCenter, and ClaimCenter) and delivers processed records as Parquet files to an AWS S3 bucket. Cloud Data Access (CDA) also maintains a manifest.json file that records the last fully committed write timestamp for each table, indicating when a batch is complete and ready to read.
The Guidewire Cloud Data Access connector is designed specifically for CDA ingestion. It reads the manifest to discover tables and their committed timestamps, then ingests Parquet files only from timestamp folders that Guidewire has marked complete. This ensures that Fivetran syncs only consistent, fully written batches to your destination.
Key functionality
The Guidewire Cloud Data Access connector provides the following key functionalities:
| Functionality | What it does | Why it matters |
|---|---|---|
| Manifest-gated ingestion | Reads the manifest before every sync and ignores any timestamp folder that Guidewire has not committed. Fivetran does not read the uncommitted folders. | Avoids ingesting uncommitted micro-batches. |
| Redeployment detection | Automatically detects new deployments and ingests their files without any manual action. Data from previous deployments is preserved. | Keeps your destination up to date across deployments. |
| Cursor-based incremental syncs | Tracks the last committed timestamp for each table and resumes from that point in subsequent syncs. | Keeps syncs efficient and bounded. |
| Schema evolution | Automatically handles data type changes, new columns, and tables based on your connection's schema change settings. | Eliminates the need for manual maintenance. |
Features
| Feature Name | Supported | Notes |
|---|---|---|
| Capture deletes | ||
| History mode | ||
| Custom data | ||
| Data blocking | ||
| Column hashing | ||
| Re-sync | ||
| Row filtering | ||
| API configurable | API configuration | |
| Priority-first sync | ||
| Fivetran data models | ||
| Private networking | ||
| Authorization via API |
Supported deployment models
We support the SaaS and Hybrid deployment models for the connector.
You must have an Enterprise or Business Critical plan to use the Hybrid Deployment model.
Setup guide
Follow our step-by-step Guidewire Cloud Data Access setup guide to connect Guidewire CDA with your destination using Fivetran.
Schema information
This schema applies to all Guidewire connections.
Sync overview
CDA file rows
Fivetran reads the committed Guidewire CDA Parquet files and loads the data into your destination without any modification. Fivetran preserves the Guidewire-specific columns such as gwcbi___operation in the destination exactly as CDA provided. Fivetran does not interpret these columns or convert them into inserts, updates, or deletes in the destination. The downstream transformation layer applies these changes.
Metadata tables
Fivetran creates the following metadata tables in your destination schema.
Manifest table
Fivetran maintains a MANIFEST table in your destination schema. During each sync, we upsert one row for each unique combination of CDA table and committed timestamp. The table has the following schema:
| Column | Data Type | Description |
|---|---|---|
table_name (Primary key) | STRING | CDA table name |
last_successful_write_timestamp (Primary key) | LONG | Timestamp of the last successful write committed for the table |
deployment_id (Primary key) | LONG | CDA deployment ID |
total_processed_records_count | LONG | Total number of records processed for the table |
schema_history | JSON | Schema history of the table as recorded in the manifest |
Batch metrics table
Fivetran maintains a BATCH_METRICS table in your destination schema. During each sync, we upsert one row for each unique combination of a CDA table, schema fingerprint, and batch timestamp from committed batch-metrics.json files. The table has the following schema:
| Column | Data Type | Description |
|---|---|---|
table_name (Primary key) | STRING | CDA table name |
schema_fingerprint (Primary key) | STRING | Schema fingerprint ID |
batch_timestamp (Primary key) | LONG | CDA batch timestamp |
deployment_id (Primary key) | LONG | CDA deployment ID |
num_records_read | LONG | Number of records CDA reads for the batch |
num_records_written | LONG | Number of records CDA writes to S3 for the batch |
num_records_dropped | LONG | Number of records CDA drops for the batch |
Deployment table
Fivetran maintains a DEPLOYMENT table in your destination schema to track the CDA deployment history. We upsert one row for each CDA deployment. We update superseded_at column value as NULL for the active deployment. The table has the following schema:
| Column | Data Type | Description |
|---|---|---|
id (Primary key) | LONG | CDA deployment epoch ID |
activated_at | TIMESTAMP | Time when the deployment was detected |
superseded_at | TIMESTAMP | Time when the deployment was replaced by a newer deployment |
Rollback table
Fivetran maintains a ROLLBACK table in your destination schema to describe the InsuranceSuite rollback events. We upsert the rollback entries present in insurancesuite-rollback-metadata.json during each sync. The table has the following schema:
| Column | Data Type | Description |
|---|---|---|
id (Primary key) | STRING | Rollback event ID |
deployment_id (Primary key) | LONG | CDA deployment ID |
lsn_lower | LONG | Lowest gwcbi___lsn value in the rollback window |
lsn_upper | LONG | Highest gwcbi___lsn value in the rollback window |
time_range_start | LONG | Beginning of the rollback window |
time_range_end | LONG | End of the rollback window |
time | TIMESTAMP | Time when the rollback entry was recorded |
Identify records affected by rollbacks
Use the following columns to identify CDC records affected by the rollback:
lsn_lower: Defines the lowestgwcbi___lsnvalue in the rollback window.lsn_upper: Defines the highestgwcbi___lsnvalue in the rollback window.time_range_start: Defines the start of the rollback window. CDC records with a timestamp earlier than this value are not affected by the rollback entry.time_range_end: Defines the end of the rollback window. CDC records with a timestamp later than this value are not affected by the rollback entry.
A CDC record is considered rolled back when its gwcbi___lsn value ranges between the lsn_lower and lsn_upper values and its gwcbi___payload_ts_ms value ranges between the time_range_start and time_range_end values. Exclude these records from your downstream queries.
Sync notes
The following notes describe how the Guidewire connector processes Parquet files and handles re-sync operations:
- The connector only ingests Parquet files from committed timestamp folders and ignores folders that Guidewire has not committed.
- It supports connection-level and table-level re-syncs. During a re-sync, Fivetran reprocesses all committed CDA timestamp folders in the active deployment.
Configuration options
You can use a custom external_id parameter for authentication when creating a new Guidewire Cloud Data Access connection with the Fivetran REST API. We use the connection's group_id if you don't specify the value of the external_id parameter. Use the List All Groups endpoint to find the connection's group_id.
We don't support custom external_id values for connections created in the Fivetran dashboard. By default, we use the connection's group_id.