As actuarial models continue to grow in complexity, the volume of data they process has increased dramatically. Whether you’re building pricing models, performing reserving analyses, running capital calculations or producing IFRS 17 results, the efficiency of your data layer has become just as important as the modelling logic itself.
For many years, CSV files have been the default way of exchanging data between systems. They’re simple, universally supported and easy to inspect with a spreadsheet or text editor. However, as datasets have grown into the millions of records, the limitations of CSV have become increasingly apparent.
That’s where Apache Parquet comes in.
Now natively supported in Mo.net (from 7.8 onwards), Parquet is a modern file format designed specifically for analytical workloads. It enables actuarial models to read data more efficiently, consume less storage and integrate seamlessly with today’s wider analytics ecosystem.
What is a Parquet File?
Unlike a CSV file, a Parquet file isn’t designed to be read by people. It’s a binary format, so opening one in a text editor simply produces unreadable characters. Instead, Parquet is designed for software.
The easiest way to understand the difference is to compare how the same data is stored.
A CSV file stores information row by row:
| PolicyID | Age | Premium | Claims |
| 1001 | 42 | 425.50 | 1 |
| 1002 | 35 | 310.25 | 0 |
| 1003 | 57 | 689.10 | 2 |
| 1004 | 29 | 275.00 | 0 |
Internally, a Parquet file stores the same information more like this:
Schema
├─ PolicyID: Integer
├─ Age: Integer
├─ Premium: Double
└─ Claims: Integer
PolicyID
1001
1002
1003
1004
Age
42
35
57
29
Premium
425.50
310.25
689.10
275.00
Claims
1
0
2
0
This isn’t the actual binary structure, but it illustrates the key concept: Parquet stores data by column rather than by row.
That difference unlocks many of the performance benefits that make Parquet so attractive for actuarial modelling.
Faster Model Execution
Consider a pricing model that only requires three variables:
- Premium
- Claim Count
- Exposure
A traditional CSV file might contain dozens or even hundreds of columns:
| PolicyID | Age | Vehicle | Region | Premium | Claim Count | Exposure | Occupation | … |
When Mo.net reads a CSV, every column has to be loaded from disk before the unwanted ones can be discarded.
With Parquet, Mo.net can read only the columns it needs.
Instead of loading the entire file, it simply retrieves:
- Premium
- Claim Count
- Exposure
while skipping everything else.
For small datasets the difference may be negligible. For experience files containing millions of policies, the reduction in disk I/O can significantly improve loading and execution times.
Smaller Files, Lower Storage Costs
Another major advantage of Parquet is file size. The format includes highly efficient compression and encoding techniques, meaning Parquet files are often 50–90% smaller than equivalent CSV files, depending on the characteristics of the data.
For actuarial teams storing years of policy history, claims experience, assumptions and model outputs, these savings can quickly become substantial.
Smaller files also mean:
- Faster backups
- Faster transfers between systems
- Lower cloud storage costs
- Quicker loading into analytical platforms
Better Data Quality Through Strong Typing
CSV files only contain text. Every application importing a CSV has to decide whether each column represents a number, date or piece of text. Regional settings, missing values and formatting differences can all introduce unexpected issues.
Parquet stores the data type alongside the data itself.
- Dates remain dates.
- Numbers remain numbers.
- Boolean values remain Boolean.
This reduces import errors and removes much of the repetitive data cleaning that often accompanies actuarial workflows.
Built for Large-Scale Analytics
Parquet has become the standard storage format across modern analytics platforms, including Python, R, DuckDB, Spark, Databricks, Microsoft Fabric and many cloud data warehouses. Because Mo.net supports Parquet directly, the same dataset can be shared across multiple analytical tools without exporting separate copies in different formats. This creates a single, consistent source of data across pricing, reserving, capital modelling and reporting teams.
Handling Large Experience Files Efficiently
Many actuarial investigations involve datasets containing tens or hundreds of millions of records. Parquet includes several features specifically designed for these workloads:
- Columnar storage minimises unnecessary reads.
- Built-in compression reduces storage requirements.
- Rich metadata allows software to skip irrelevant sections of a file.
- Partitioning makes it easy to work with individual years, products or business units.
Together, these capabilities allow Mo.net to process large datasets far more efficiently than traditional text-based formats.
A Better Foundation for Automated Data Pipelines
Modern actuarial modelling increasingly forms part of automated production processes. Data flows from operational systems into data lakes, through validation pipelines, into actuarial models and finally into reporting platforms.
Parquet fits naturally within these workflows because it preserves both data structure and data types. The result is more reliable automated processes, fewer manual interventions and reduced risk of data formatting issues disrupting production runs.
Why It Matters for Mo.net Users
Support for Parquet within Mo.net allows actuarial teams to benefit from industry-standard data technology without changing the way they build models.
Users can take advantage of:
- Faster data loading
- Reduced model execution times
- Smaller data files
- Lower memory consumption
- More reliable handling of data types
- Better integration with modern analytics platforms
- Improved scalability as datasets continue to grow
These benefits apply across pricing, reserving, capital modelling, IFRS 17 and experience investigations.
Looking Ahead
The actuarial profession is rapidly embracing larger datasets, cloud-native architectures and increasingly sophisticated analytical techniques. Parquet has emerged as one of the foundational technologies enabling that transition.
By supporting Parquet, Mo.net gives actuarial teams access to a faster, more scalable and more robust way of managing model data. Rather than treating data loading as a bottleneck, modellers can focus on what matters most—developing better models and delivering insights more quickly.
For organisations looking to modernise their actuarial workflows, adopting Parquet isn’t simply a technical upgrade. It’s an investment in a data platform that’s designed for the future.