Technical
Technical FAQ
How Aethora Labs Inc. extracts, transforms, de-identifies, secures and disposes of data across its lifecycle. Every answer is mapped to the control or framework it derives from — ISO/IEC 27001, ISO/IEC 27701, the NIST Privacy Framework and NIST AI RMF — and the final section covers the Cyber and Data Risk insurance that sits behind those controls.
Effective 28 August 2026
The core architectural principle
Separate intelligence from identity as early in the data lifecycle as technically feasible. The objective is to retain the operational relationships AI systems need — decisions, context, exceptions, sequences and outcomes — while minimizing information that identifies the individuals behind those events.
Assessment & extraction scope
| Question | Response | Basis |
|---|---|---|
| How is source data initially assessed? | Before extraction, Aethora performs data discovery and classification to identify source systems, file types, schemas, structured/unstructured content, sensitive fields, direct identifiers, quasi-identifiers, and the operational variables relevant to the approved use case. | ISO 27001 / ISO 27701 / NIST Privacy Framework |
| How is extraction scope determined? | Extraction follows a purpose-defined schema rather than indiscriminate collection. Required tables, fields, documents, date ranges and relationships are identified before processing. Fields without a defined intelligence purpose are excluded where technically practicable. | Data Minimization / Purpose Limitation |
| Do you replicate the client's entire database? | Not by default. Aethora favors selective extraction of approved datasets and fields over unrestricted replication of production environments. | ISO 27701 / NIST Privacy Framework |
| How are databases extracted? | Depending on the environment, extraction may occur through approved read-only database connections, APIs, exports, ETL/ELT pipelines, secure file transfer, or controlled client-generated extracts. Methods are selected according to system architecture and risk. | ISO 27001 / Secure Processing |
| Do you require write access to production systems? | Extraction workflows should use read-only or least-privileged access wherever technically feasible. Production write privileges are not inherently required for intelligence extraction. | ISO 27001 / Least Privilege |
| How are extraction credentials controlled? | Service accounts and credentials are scoped to the minimum permissions necessary. Credentials should be centrally managed, protected from source code, rotated according to policy, and revoked when no longer required. | ISO 27001 / Access Control |
| How do you handle APIs? | API-based extraction is performed through authenticated endpoints with scoped permissions. Where supported, Aethora can apply rate limiting, restricted scopes, short-lived credentials, IP/network restrictions and auditable service identities. | ISO 27001 |
De-identification & privacy engineering
| Question | Response | Basis |
|---|---|---|
| How do you prevent unnecessary PII from entering the pipeline? | Where technically feasible, field-level filtering occurs at or near the source so unnecessary identifiers are never included in the working extraction. Where source-side exclusion is impractical, identifiers are isolated and transformed at the earliest controlled processing stage. | Privacy by Design / ISO 27701 |
| How are direct identifiers detected? | Detection can combine schema inspection, field-name classification, pattern matching, data-type analysis, metadata review, and content-level detection for structured and unstructured data. Identified fields are mapped to the applicable treatment rule before downstream packaging. | NIST Privacy / ISO 27701 |
| How do you handle unstructured documents? | Documents can undergo parsing and content classification to identify operational content and sensitive information. Relevant intelligence is then extracted while unnecessary identifying content is removed, redacted, transformed or excluded according to the approved processing rules. | NIST Privacy Framework |
| Is simply deleting names considered de-identification? | No. Aethora distinguishes direct identifiers from quasi-identifiers and contextual identifiers. Removing a name does not by itself eliminate identification risk where other attributes could reasonably be linked back to an individual. | NIST De-identification Guidance |
| How are identifiers transformed? | Depending on the use case, treatment may include suppression, redaction, masking, tokenization, pseudonymization, generalization, aggregation, hashing where appropriate, or removal. The technique is selected according to the required analytical utility and identification risk. | NIST / ISO 27701 |
| Do you treat hashing as anonymization? | Not automatically. A deterministic hash of predictable information may remain susceptible to matching or dictionary attacks. Where pseudonymous linkage is necessary, stronger keyed transformations or tokenization may be appropriate, with keys managed separately from the resulting dataset. | Privacy Engineering / NIST |
| Can records remain linkable after identity removal? | Where the approved use case requires longitudinal analysis, records may be assigned non-identifying surrogate or pseudonymous identifiers that preserve relationships across events without exposing the underlying direct identity. | ISO 27701 / Privacy Engineering |
| How are tokenization keys protected? | Where reversible tokenization or keyed transformations are used, key material should be segregated from transformed datasets, access-restricted, centrally managed, and protected according to cryptographic key-management procedures. | ISO 27001 |
| How is re-identification risk evaluated? | Aethora evaluates whether combinations of retained attributes could materially increase identification risk. Mitigation can include generalization, bucketing, suppression, aggregation, rare-value treatment or additional access restrictions. | NIST De-identification Guidance |
| How do you preserve intelligence after de-identification? | The objective is utility-preserving transformation: remove identity while retaining relationships necessary to understand decisions, sequences, exceptions, classifications, interventions and outcomes. Validation occurs against the defined intelligence use case rather than simply counting remaining fields. | Privacy Engineering / Data Quality |
Security, access & encryption
| Question | Response | Basis |
|---|---|---|
| How is data encrypted in transit? | Transfers should use modern encrypted transport protocols such as TLS-protected APIs or approved secure file-transfer mechanisms, according to the environment and client requirements. | ISO 27001 |
| How is data encrypted at rest? | Sensitive datasets should be stored using encryption-at-rest capabilities appropriate to the underlying storage environment, with cryptographic keys managed separately from ordinary user access where supported. | ISO 27001 |
| How is access segmented? | Aethora applies role-based and least-privileged access so extraction, engineering, review, administrative and delivery functions do not automatically receive identical access to source and transformed information. | ISO 27001 |
| Is administrative activity logged? | Security-relevant activity should be logged, including authentication events, privileged actions, access to sensitive environments, material configuration changes and relevant data-processing events. | ISO 27001 / NIST |
| Can you determine who accessed a dataset? | Controlled environments should associate access with identifiable users or service identities, allowing authorized activity to be traced through appropriate audit logs. Shared credentials should be avoided for sensitive processing. | ISO 27001 |
| Are raw and transformed datasets separated? | Yes, where the architecture requires retention of both. Raw ingestion, transformation workloads and approved output datasets should be logically or physically segregated with access determined independently for each processing zone. | ISO 27001 / Defense in Depth |
Integrity, lineage & versioning
| Question | Response | Basis |
|---|---|---|
| How do you maintain data lineage? | Processed datasets should maintain traceable metadata identifying the source, extraction event, transformation stage, applicable processing rules, dataset/version identifier and resulting output — establishing provenance from source through licensed artifact. | ISO 27001 / NIST AI RMF |
| How do you prevent the raw dataset from becoming the licensed product? | Aethora's pipeline distinguishes source data from derived intelligence artifacts. Only datasets that have completed the required transformation, privacy, quality and authorization gates are eligible for downstream packaging or delivery. | Privacy by Design / Data Governance |
| How do you validate extraction completeness? | Extraction can be validated through record counts, schema validation, expected-field checks, reconciliation, null/error analysis, and checksums or hashes where appropriate, alongside source-to-output quality controls. | Data Quality / Processing Integrity |
| How do you detect pipeline corruption or unauthorized modification? | Depending on the architecture, integrity controls can include cryptographic hashes, immutable/versioned storage, audit logging, reconciliation checks, controlled deployment processes and source-to-output validation. | ISO 27001 / Processing Integrity |
| How are dataset versions controlled? | Material transformations should produce identifiable dataset versions so Aethora can distinguish source snapshots, transformation logic, corrected datasets, refreshes and approved releases. | Data Lineage / NIST AI RMF |
| How do you handle dataset refreshes? | Refreshes are treated as controlled processing events. New source material is extracted according to the approved scope, processed through applicable transformation rules, validated, versioned and authorized before incorporation. | ISO 27001 / NIST AI RMF |
AI-training data preparation & permitted use
| Question | Response | Basis |
|---|---|---|
| How is AI-training data prepared? | Depending on the use case, preparation may include normalization, deduplication, classification, structuring, metadata enrichment, segmentation, quality filtering, transformation and provenance tagging before the dataset is approved for the intended AI application. | NIST AI RMF / Data Governance |
| How do you prevent duplicate data from distorting the dataset? | Deduplication can use exact identifiers, hashes, normalized field comparisons or similarity-based techniques appropriate to the data type, while maintaining sufficient lineage to understand why records were consolidated or excluded. | Data Quality / AI Data Governance |
| How do you document permitted downstream use? | Each dataset should have a defined purpose, authorized use, provenance, applicable restrictions, processing history and licensing scope. Technical access controls and contractual controls operate together rather than relying exclusively on either. | ISO 27701 / NIST AI RMF |
| Can a buyer access the original identifiable source data? | Aethora's model is designed around licensing the approved intelligence artifact, not providing unrestricted access to the client's underlying identifiable source environment. Any exception would require separate authorization and applicable controls. | Privacy by Design / Purpose Limitation |
Retention, disposal & incidents
| Question | Response | Basis |
|---|---|---|
| How are temporary processing files handled? | Temporary exports, intermediate files, caches and staging artifacts are governed by the same classification principles as their source data and removed according to defined processing and retention procedures. | ISO 27001 / ISO 27701 |
| What happens when processing is complete? | Source and intermediate data are retained only according to the applicable contractual, operational, security or legal requirements. Data no longer required is subject to established retention and secure-disposition procedures. | ISO 27001 / ISO 27701 |
| How is deletion verified? | Where required by the engagement, disposition can be documented through workflow records, system evidence, deletion logs or other appropriate evidence based on the storage architecture involved. | ISO 27001 |
| How do you handle security incidents involving client data? | Suspected incidents are handled through a defined incident-management process covering identification, containment, investigation, remediation, documentation, escalation and contractual or legally required notification. | ISO 27001 / NIST |
Cyber & data risk insurance
Insurance operates as a residual risk-transfer mechanism. It does not replace Aethora's primary controls — data minimization, least-privileged extraction, encryption, segregation, de-identification, access control, logging, lineage, integrity validation, retention controls and incident response — it provides an additional financial layer if preventive and detective controls do not fully contain an event.
| Question | Response |
|---|---|
| Does Aethora maintain cybersecurity insurance? | Yes. Aethora Labs maintains Cyber and Data Risk insurance through Hiscox Insurance Company Inc., providing an additional financial risk-transfer layer alongside Aethora's technical, administrative and contractual security controls. |
| What is the current coverage limit? | Aethora's current policy provides $1,000,000 per claim and $1,000,000 aggregate coverage. |
| Does the policy address data breaches? | Yes. Breach-response coverage includes forensic investigation, affected-individual notification, credit-protection services, call-center costs, crisis management and public-relations expenses. |
| Does it address privacy liability associated with handling client data? | Yes. Privacy Protection covers defense and resolution of claims concerning personally identifiable or confidential corporate information, including negligence, privacy or consumer-protection violations, breach of contract, regulatory investigations and certain network-security failures. |
| What about ransomware or cyber extortion? | The policy includes cyber-extortion coverage for response costs and financial payments associated with network-based ransom demands, subject to the applicable policy terms and ransomware sublimit. |
| What happens if data or software is damaged during a cyber event? | Data Recovery covers costs associated with replacing, restoring or repairing damaged or destroyed data and software, subject to policy terms. |
See also our Governance & Compliance program and our Privacy Policy.
