Crayon Data Crayon Data Tangram AI
Data onboarding & enrichment | Fully agentic |
DATA · ONBOARDING AND ENRICHMENT SUITE
DATA STUDIO

Ingests, validates, enriches and governs data from every source you connect — end to end
so a clean, tagged, compliance-ready dataset lands where your analysts and agents can actually use it.

70% less prep time 30% better governance 80% of metadata automated
Why Data Studio exists

Weeks of data prep. Before any value is created.

The problem

Data arrives in every shape

  • Every sourcedatabases, buckets, APIs, file drops
  • No two schemasagree on a field name
  • Silent errorsnulls, duplicates and drift pass through
Today, by hand

Scripts, tickets and spreadsheets

  • 50–70%of analyst time spent preparing, not analysing
  • Per sourcemapping rules written and maintained by hand
  • ManualPII discovery and metadata tagging
With Data Studio

Preparation stops being the project

  • 70%less data preparation time
  • 30%higher quality and governance efficiency
  • 80%of metadata work automated
What it does · three moves

Land it. Trust it. Govern it.

01Ingest & MapLand it in one model
  • Any source — database, bucket, API or file drop
  • Deduplicated and normalised on arrival
  • Fields mapped by context, no ruleset written
02Validate & EnrichMake it trustworthy
  • AI checks quality and flags anomalies
  • Categories, attributes and context added
  • Bad records held, never silently passed
03Govern & ProtectLabel it before it moves
  • PII, PCI and sensitivity tagged automatically
  • Masking and role-based access applied
  • Every tag change on the audit log
MOVE 01Ingest from anywhere, land in one model.CONNECTED SOURCESPostgreSQLJDBC · 4 schemasS3 BucketPARQUET · dailyREST APIOAUTH2 · hourlyKafka streamAVRO · continuousCSV over SFTPON ARRIVALINGEST1.28Mrows normalisedDEDUPED ON KEYSCHEMA ALIGNMENTcust_nocustomer_idmrchnt_nmmerchant_nametxn_amtamountdt_stampevent_timeccycurrencyMAPPED BY CONTEXT · NO RULESETANY SOURCE · NO CODE WRITTEN
MOVE 02Catch what is wrong, add what is missing.QUALITY GATE96.4%passed all rules449 HELD FOR REVIEWRULES RUN THIS PASSNull check · customer_idPASSRange · txn_amountPASSFormat · iso_timestampFLAGAnomaly · velocity spikeFLAGDuplicate · composite keyPASSENRICHMENT APPLIEDMERCHANT_NAMENordstromCATEGORYDepartment StoreADDEDMCC_CODE5651 · ApparelADDEDBRAND_LOGOcdn/nordstrom.svgADDEDGEOSeattle, WAADDEDCONTEXT ADDED · SOURCE UNTOUCHED
MOVE 03Sensitive data is labelled before it moves.COLUMN CLASSIFICATIONcustomer_panPII · RESTRICTEDmask on reademail_addressPII · RESTRICTEDmask on readcard_numberPCI · SECRETtokenisedtxn_amountINTERNALrole-gatedmerchant_namePUBLICopenmcc_codePUBLICopenHELD AGAINSTGDPRArt. 9 special category mappedPCI DSSRequirement 3.4 satisfiedCCPAOpt-out propagated downstreamAudit logEvery tag change recordedAPPLIED BEFORE ANYTHING LEAVES THE PLATFORM
The capability set

Six things it does, on every dataset.

Not a menu of optional modules — all six run on every source you connect.

01Data ingestionExternal sources, collected automatically
02Data enrichmentRaw records gain useful context
03Metadata managementOrganised and maintained for you
04Schema mappingStructures aligned across sources
05Sensitivity taggingSensitive fields labelled for compliance
06AI-driven validationQuality checked, anomalies detected
S1 → S7

Seven stages. One clean dataset.

S1Source Configuration

User configures external data sources to connect.

S2Data Ingestion

System automatically pulls data from configured sources.

S3Validation ProcessingQuality gate

AI validates data quality and flags issues.

S4Enrichment Application

Adds contextual information to raw datasets.

S5Metadata Tagging

Applies sensitivity and classification tags automatically.

S6Schema Alignment

Maps and aligns data structures across sources.

S7Delivery

Clean, enriched dataset ready for analysis.

S1Source ConfigurationUser configures external data sources to connectPostgreSQLJDBC · 4 schemasS3 BucketPARQUET · dailyREST APIOAUTH2 · hourlyCSV DropSFTP · on arrivalDATA STUDIO4sources connectedCREDENTIALS ENCRYPTEDNO CODE WRITTENS1S2S3S4S5S6S7
S2Data IngestionSystem automatically pulls data from configured sourcesTXN_IDMERCHANTAMOUNTTSTX-88412NORDSTROM1,240.0010:04:11TX-88413SHELL62.4010:04:12TX-88414AMAZON318.9010:04:12TX-88415STARBUCKS7.2510:04:13TX-88416DELTA AIR902.0010:04:13INGESTED1,284,902rows this runBATCHstreaming + batchDEDUPEon primary keyERRORS0 droppedS1S2S3S4S5S6S7
S3Validation ProcessingAI validates data quality and flags issuesRULE RESULTSNull check · customer_id1,284,902 / 1,284,902PASSRange · txn_amountwithin 0 – 50,000PASSFormat · iso_timestamp412 malformedFLAGReferential · merchant_idno orphansPASSAnomaly · velocity spike37 accountsFLAGDuplicate · composite key0 duplicatesPASSQUALITY SCORE96.4%449 records flaggedHELD FOR REVIEWS1S2S3S4S5S6S7
S4Enrichment ApplicationAdds contextual information to raw datasetsRAW RECORDMERCHANT_IDM-40912RAW_NAMENORDSTROM #418AMOUNT1240.00CURRENCYUSDENRICHENRICHED RECORDMERCHANT_NAMENordstromCATEGORYDepartment StoreADDEDMCC_CODE5651 · ApparelADDEDBRAND_LOGOcdn/nordstrom.svgADDEDGEOSeattle, WA · 47.61,-122.33ADDEDS1S2S3S4S5S6S7
S5Metadata TaggingApplies sensitivity and classification tags automaticallyCOLUMN CLASSIFICATIONcustomer_panPII · RESTRICTEDmask on reademail_addressPII · RESTRICTEDmask on readcard_numberPCI · SECRETtokenisedtxn_amountINTERNALrole-gatedmerchant_namePUBLICopenmcc_codePUBLICopenCOMPLIANCEGDPRArt. 9 mappedPCI DSSreq. 3.4 metCCPAopt-out honouredAUTO-APPLIEDS1S2S3S4S5S6S7
S6Schema AlignmentMaps and aligns data structures across sourcesSOURCE SCHEMATARGET MODELcust_nomrchnt_nmtxn_amtdt_stampccycustomer_idmerchant_nameamountevent_timecurrency99%97%99%94%99%INFERRED BY CONTEXT · NO RULESET WRITTENS1S2S3S4S5S6S7
S7DeliveryClean, enriched dataset ready for analysisCUSTOMER_IDMERCHANT_NAMECATEGORYAMOUNTC-10041NordstromDept. Store1,240.00C-10042ShellFuel62.40C-10043AmazonE-commerce318.90C-10044StarbucksDining7.25C-10045Delta AirTravel902.00C-10046Whole FoodsGrocery146.10DELIVERED TOAnalytics warehouseSNOWFLAKEFeature storeML PIPELINESBI layerDASHBOARDSDownstream agentsGOVERNED ACCESSLINEAGE SEALED · FULLY AUDITABLES1S2S3S4S5S6S7
Key benefits · agent vs. manual

The same work, at a fraction of the effort.

Preparation time

Reduces data preparation by 50–70%
0%Less time from raw feed to usable dataset

Data quality

Fewer defects reaching downstream systems
0%Higher quality and governance efficiency

Metadata management

Manual tagging effort largely removed
0%Of metadata work handled automatically
Agents and models powering it

Nine agents, working as one.

Data Studio is a bundle. Each agent owns a stage of the pipeline — you deploy the suite, not nine integrations.

01Anomaly InvestigatorFraud, errors and unexpected activity, in real time
02Natural Language QueryPlain-English questions, governed answers with citations
03Data CleansingDeduplicates and validates for ML readiness
04Data MappingMaps fields across schemas without a ruleset
05Data QualityRules declared in English, SQL checks generated
06Data MaskingObfuscates PII, financial and health data in real time
07ValidationChecks dynamic merchant and offer fields against rules
08Metadata EnrichmentAssigns attributes, tags, URLs and images
09Transaction OnboardingNormalises high-volume transaction feeds
Tech · security · governance

Enterprise-grade, from the foundation up.

Core platformFOUNDATIONFastAPI · SQLAlchemy · Pydantic · PostgreSQL · MySQL · UltraDB
AI & automationINTELLIGENCEOpenAI GPT-4 · LangChain · natural language to SQL
SecurityCONTROLSEncrypted credentials · SSL/TLS · audit logging · role-based access
DeploymentWHERE IT RUNSOn-premise · Docker · hybrid · container-ready architecture
In summary

Clean, enriched, governed data — ready to use.

0%Less preparation time
0%Better governance efficiency
0%Of metadata automated
0 agentsBundled in the suite
See it liveBook a working demoTry itHosted sandbox · your own sample dataDeployOn-premise · Docker · hybrid
Data Studio · Crayon Data

Thank you.

Scan for the live portal

QR code to the live portal