Automating ISO 45001 Hazard Identification in Jurong Island Refineries Using LLM-Powered Document Analysis

Introduction

LLM-powered document analysis can automate 70–80% of manual hazard identification tasks in refinery safety documentation, reducing analysis time from weeks to hours while strengthening ISO 45001 compliance across Jurong Island’s petrochemical operations. For HSE managers, safety engineers, and compliance officers operating within one of the world’s top three oil refining centres, this shift from manual document review to intelligent automation represents both a practical efficiency gain and a fundamental improvement in how potential hazards are detected, classified, and managed.

This article focuses specifically on automating hazard identification for refinery operations on Jurong Island-Singapore’s core energy and chemicals hub where over 100 companies in petrochemicals, process utilities, and storage operate in close proximity. It does not cover general industrial artificial intelligence applications or non-refinery safety contexts. The target audience comprises process safety professionals, HSE leadership, and process engineer teams responsible for maintaining ISO 45001 occupational health and safety management systems in major hazard installations. Hazard identification is essential for effective risk mitigation, and oil and gas facilities lose $8.4M annually to safety incidents-making the business case for automation compelling.

Large Language Models can enhance the identification of hazards in complex environments like refineries by processing unstructured safety documentation-incident reports, operating procedures, piping and instrumentation diagram descriptions, maintenance logs-and extracting structured hazard data at scale. Hazard identification precedes all safety activities in process safety, and automating this foundational step yields cascading benefits across the entire risk management lifecycle.

By the end of this article, you will understand:

  • How LLM-powered document analysis detects and classifies safety hazards from refinery documentation

  • A practical 8-week implementation framework tailored to Jurong Island facilities

  • Quantified performance benchmarks from real-world case studies in process safety

  • Strategies for managing false positives, document quality issues, and regulatory compliance

  • How automated hazard identification integrates with existing risk assessment methodologies and digital technologies

An aerial view of a petrochemical industrial island showcases interconnected refinery units, large storage tanks, and various processing facilities, all surrounded by water. This image highlights the complexity of operations involved in hazard identification and risk assessment processes within the industry.

Understanding ISO 45001 Hazard Identification and Process Hazard Analysis in Refinery Operations

ISO 45001:2018 defines hazard identification as the systematic process of recognizing sources or situations with the potential to cause injury or ill health. Under Clause 6.1.2, organizations must proactively identify hazards covering routine and non routine activities, human factors, emergency situations, design modifications, and all persons who may be affected-including workers, contractors and visitors. ISO 45001 requires documented hazard identification procedures, and this requirement extends to ongoing hazard identification and risk assessment as operational conditions change. For refineries, this means cataloguing an enormous range of hazard types: chemical exposure (hydrogen sulfide, benzene), fire and explosion risks from hydrocarbon release, mechanical failure mode scenarios in rotating equipment, thermal hazards, confined space dangers, and electrical risks.

Current manual processes used in Jurong Island refineries for document review involve cross-disciplinary teams of process engineers, mechanical engineers, and HSE professionals systematically reviewing process flow diagrams, Piping & Instrumentation Diagrams (P&IDs), Standard Operating Procedures (SOPs), incident and near-miss reports, maintenance logs, and Material Safety Data Sheets. Refineries use different methods for hazard identification depending on the study objective and project phase, with techniques such as hazard and operability study (HAZOP), process hazard analysis, What-If analysis, and mechanical integrity reviews producing hazard recognition, risk assessment findings, and recommended controls in traditional risk assessment methodologies. Protection analysis is then often used to test whether existing safety layers are adequate after hazards are identified. These processes are time consuming, often spanning weeks to months per process unit, and are subject to human error, fatigue, and inconsistent terminology across documents.

Traditional Documentation Review Challenges

The volume and complexity of refinery safety documentation present significant barriers to thorough hazard identification. A single process unit generates thousands of pages across design documentation, SOPs, inspection records, and maintenance histories. P&IDs alone for a crude distillation unit can span hundreds of sheets, each containing critical information about equipment, control loop configurations, safety integrity level specifications, and emergency shutdown systems. Inadequate hazard identification led to four fatalities at a refinery-a stark reminder that missed hazards in these documents carry lethal consequences.

Manual review demands subject-matter experts with deep process knowledge, and the time-intensive nature of the work creates bottlenecks. Human error potential increases with document volume: varied terminology (synonyms like “flare tip,” “vent tip,” and “blow-down” referring to the same equipment), inconsistencies across documents from different eras, and legacy formats including scanned PDFs and handwritten annotations all degrade review quality. On Jurong Island specifically, the density and adjacency of facilities means that associated risks may arise from inter-unit or inter-company interactions-domino effects that are rarely captured in a single unit’s documentation review. Hazard identification involves transforming unstructured data into structured insights, and manual processes struggle to achieve this transformation consistently at scale.

A safety professional is seated at a desk, meticulously reviewing large stacks of technical drawings, inspection reports, and procedure manuals, all featuring highlighted annotations. This scene reflects the critical process of hazard identification and risk assessment, essential for ensuring safety in operations and managing associated risks in the workplace.

Regulatory Requirements for Jurong Island Facilities

Jurong Island refineries must comply with Singapore’s Workplace Safety and Health Act (WSHA), which places clear duties on employers and occupiers to ensure safety, investigate dangerous occurrences, and enforce compliance. Refineries are Major Hazard Installations regulated by the Major Hazards Department in Singapore, meaning hazard assessment must also consider off-site consequences, community exposure from potential hazards, and proximity to neighbouring facilities. Effective hazard identification is mandatory under OSHA PSM regulations, and Singapore’s parallel framework under SS 506 Part III (“Code of Practice for Process Safety”) imposes additional obligations for highly hazardous installations on Jurong Island.

Industry-specific hazard categories relevant to refinery operations span process safety incidents (fire, explosion, toxic release), occupational health exposure (chronic chemical exposure, noise, ergonomic harm), confined space work activities, high-temperature operations, and mechanical integrity failures. The Jurong Island Vision Zero safety cluster-comprising major players like ExxonMobil, Shell, and PCS-emphasizes proactive hazard identification, behaviour-based safety, contractor integration, and process safety systems. Automated risk assessments require strict adherence to local regulations and safety standards, making any automated system’s alignment with both ISO 45001 and WSH requirements non-negotiable.

Continuous monitoring of operational changes is essential to maintaining safety standards, and this regulatory expectation creates a natural fit for automated systems that can process new documents and flag emerging risks in near real time monitoring configurations.

LLM and Machine Learning-Powered Document Analysis Technology

Large Language Models-transformer-based artificial intelligence systems trained on massive text corpora-bring powerful natural language processing capabilities to technical documentation analysis. When applied to refinery hazard identification, these models can parse unstructured text from incident reports, maintenance records, and operating procedures to extract hazard entities, classify risk levels, map causal chains linking immediate causes to organizational factors, and generate structured outputs aligned with ISO 45001 categories. The technology bridges the gap between the overwhelming volume of hse data in refinery operations and the structured hazard registers required for compliance and effective risk management.

The practical application in refinery hazard identification workflows builds on demonstrated successes across process safety research. Natural language processing automates hazard identification from design documents, and several architectures-from fine-tuned encoder models like BERT to few-shot prompted decoder models like GPT-4o-have shown strong performance in extracting safety-critical information from industrial texts. A strong configuration for hazard identification involves integrating site-specific data with LLMs, ensuring that the model understands local terminology, equipment naming conventions, and facility-specific risk profiles.

Document Types and Analysis Capabilities

The primary input sources for LLM-powered hazard analysis include technical specifications, incident and accident reports, maintenance records, inspection logs, operational procedures, design basis documents, regulatory submissions, and descriptions extracted from instrumentation diagram documentation. Each document type contributes different facets to the hazard identification picture: incident reports reveal actual failure mode patterns and identified risks; SOPs expose routine maintenance and non routine task hazards; P&ID descriptions identify equipment configurations, safety devices, and control loop arrangements; and design documentation captures original safety consideration decisions including safety integrity level assignments.

Research demonstrates compelling accuracy benchmarks. A BERT-based risk feature identification model applied to an oil refinery context-embedded in a web application called HALO-achieved approximately 97.42% accuracy for consequence prediction, 86.44% for severity categorization, and 94.34% for likelihood estimation. A chemical entity recognition model using BERT-BiLSTM-Self-Attention-CRF architecture achieved F1 scores of approximately 94.57% for recognizing hazardous chemical risk entities and relationships. Auto-BIMHazard achieves 84.77% accuracy in hazard classification using machine learning approaches, and a 1D-CNN model identifies hazards with 71% accuracy in construction contexts-suggesting that refinery-specific fine-tuning can push performance even higher.

For causal chain extraction, a study processing chemical accident reports achieved entity extraction F1 of approximately 0.77 and causal triplet extraction F1 of approximately 0.73 using chain-of-thought prompting with GPT-4o. This performance approached fine-tuned BERT-large levels, and calibrated confidence scoring allowed routing approximately 87% of errors to human experts via low-confidence flagging-dramatically reducing expert review workload. Retrieval-Augmented Generation can enhance the context and accuracy of hazard identification outputs by grounding model responses in facility-specific documentation rather than relying solely on pre-trained knowledge.

The image displays a split screen, with a traditional text document on the left side and the same document on the right, enhanced with AI-powered color-coded hazard annotations, risk classifications, and causal chain diagrams. This visual representation aids in hazard identification and risk assessment, showcasing the integration of digital technologies in the risk management process for improved safety in industrial operations.

Hazard Classification, Categorization, and Risk Assessment Methodologies

Automated sorting of identified hazards into ISO 45001-compliant categories enables direct population of hazard registers and risk assessment matrices. LLMs can classify extracted hazards into standard categories-physical, chemical, biological, ergonomic, psychosocial-as well as process safety-specific categories including fire, explosion, toxic release, and process equipment failure. The system maps each identified hazard to severity and likelihood scores, generating quantitative risk assessment outputs that align with existing refinery risk matrices.

Integration with existing refinery hazard registers allows the system to compare newly identified hazards against previously identified risks, flagging novel findings and gaps in coverage. This is particularly valuable for identifying low-frequency but high-consequence hazards that may be buried in legacy documents or obscured by inconsistent terminology. HAZID studies identify major hazard categories early in projects, and automated classification can replicate this categorization across thousands of documents simultaneously, applying consistent evaluation criteria that eliminate the variability inherent in manual effects analysis.

The predictive capability of these systems extends beyond simple extraction. Integration of historical data can reveal recurring hazard patterns and improve safety practices, and LLMs can synthesize patterns across years of incident reports to identify systemic risks that individual reviewers might miss. AI detects safety risk precursors in real time, and machine learning improves near-miss detection by 68–84%-capabilities that transform hazard identification from a periodic compliance exercise into a continuous safety improvement process.

Implementation Framework for Jurong Island Refineries

Deploying LLM-powered hazard identification in Jurong Island refineries requires a structured approach that accounts for the unique characteristics of major hazard installations: stringent regulatory requirements, proprietary documentation, complex process configurations, and the need for absolute reliability in safety-critical applications. The following framework builds on demonstrated capabilities from process safety research while addressing Jurong Island’s specific operational context. Automated hazard identification can be integrated into existing systems in under 2 weeks for initial setup, though a comprehensive deployment benefits from the full 8-week methodology described below.

8-Week Deployment Methodology

A phased implementation approach minimizes operational disruption while building confidence through validated results at each stage.

Weeks 1–2: Document Inventory and System Integration Planning. Conduct a comprehensive audit of all relevant documents across process units-SOPs, design documentation, P&IDs, incident reports, maintenance logs, and inspection records. Evaluate document formats, quality (identifying OCR needs for scanned materials), metadata completeness, and naming conventions. Map integration points with existing document management systems, SCADA/DCS platforms, and work-order systems. Establish the hazard taxonomy aligned with both ISO 45001 categories and Singapore WSH process safety categories.

Weeks 3–4: LLM Training on Refinery-Specific Hazard Terminology and Documentation Formats. Select base models appropriate to the task-fine-tuned BERT variants for entity extraction, domain-adapted LLMs for causal reasoning and classification. Annotate sample documents from local plants with hazard entity labels, severity and likelihood classifications, and causal relationships. Develop prompt protocols for few-shot or zero-shot scenarios. Define confidence thresholds and build the local glossary mapping synonyms and company-specific acronyms to standardized terms. Storage and processing of sensitive safety information must comply with data governance standards, so establish data handling protocols during this phase.

Weeks 5–6: Pilot Testing on Selected Process Units and Validation Against Manual Reviews. Deploy the model on documents from one to two well-characterized process units-crude distillation or catalytic reforming units where mature hazard registers exist. Compare automated outputs against manual hazard registers: measure the number of hazards identified, novel findings, misses, and false positives. Quantify processing time versus manual review duration. Hazard identification should include engagement with certified safety professionals for validation, and expert review during this phase calibrates model thresholds and refines taxonomy.

Weeks 7–8: Full System Deployment with Continuous Monitoring and Refinement. Roll out across the full facility or multiple process units. Integrate with document management and reporting systems for near-real-time processing of new incident reports and maintenance logs. Configure dashboards feeding into HSE audits, hazard registers, and compliance reporting. Establish feedback loops where manual corrections feed back into model fine-tuning. Define ongoing performance KPIs.

The image depicts a visual timeline outlining four two-week phases of an industrial technology deployment, highlighting milestone markers, deliverables, and team activities at each stage. This structured approach aids in hazard identification and risk assessment processes, ensuring comprehensive management of associated risks throughout the deployment.

System Integration and Data Sources

Integration with existing document management systems ensures that the LLM processes the full corpus of available safety documentation without manual upload requirements. Connecting to SCADA and DCS platforms enables the system to cross-reference process data with document-based hazard identification-for example, correlating emergency shutdown events logged in SCADA with the hazard analysis documented in SOPs. AI predicts H2S events 25–60 minutes before alarms trigger, demonstrating the value of integrating sensor data with document-based hazard intelligence. Real-time processing capabilities mean that new incident reports, maintenance findings, and inspection records are automatically analysed as they enter the system-a critical advantage over periodic manual reviews that may leave gaps of weeks or months. Data quality and provenance are critical for safety assessments in industrial contexts, and the integration layer must maintain full traceability of source documents, extraction timestamps, and model versions.

Performance Metrics and Validation

Key accuracy benchmarks for comparing automated versus manual hazard identification results include precision, recall, and F1-score for hazard entity extraction; accuracy for severity and likelihood classification; and causal chain extraction completeness. Based on published research, realistic targets include entity extraction F1 of 0.77 or above, severity classification accuracy exceeding 86%, and consequence prediction accuracy approaching 97%-benchmarks established by the HALO model and chemical accident report extraction studies.

Processing speed improvements are substantial: manual hazard identification for a new process unit typically requires 100+ engineering person-hours spanning several weeks. Automating 70–80% of this workload reduces review time to days, with expert validation focused on the subset of low-confidence or high-severity findings. Predictive hazard identification can reduce incident costs by $520,000 per facility annually, and automated systems can reduce near-miss incidents by 84%. Automation can improve the efficiency of incident reporting and hazard identification processes across the board, while automated systems can assist in generating Job Safety Analyses from historical incident data-further multiplying the productivity gains.

Common Implementation Challenges and Solutions

Deploying LLM-powered hazard identification in active refinery environments introduces challenges that differ from typical enterprise AI implementations. The safety-critical nature of the application, combined with regulatory scrutiny and proprietary documentation concerns, demands careful planning across several dimensions. LLMs should support but not replace human decision-making in safety-critical environments-a principle that must guide every design decision.

Document Quality and Standardization Issues

Legacy documents in older Jurong Island plants may include design drawings from multiple eras, scanned PDFs with degraded image quality, handwritten annotations, and contractor documents that lack standard formatting or metadata. These document quality issues directly degrade model performance by introducing OCR errors, missing context, and inconsistent terminology.

Solution: Implement robust preprocessing protocols that include high-quality OCR with confidence scoring, format normalization pipelines, and standardized templates for new documentation. Develop plant-level hazard glossaries that map synonyms, abbreviations, and company-specific acronyms to standardized terms. For ongoing operations, enforce template usage for all new documentation-incident reports, routine maintenance logs, inspection records-to ensure consistent machine readability. CHPtD enhances safety by embedding hazard identification in design, and similarly, embedding documentation standards at the creation point prevents downstream quality issues.

The image shows a side-by-side comparison of an industrial document; on the left, a faded scanned version, and on the right, the same document after digital processing, featuring clean text, structured data fields, and standardized formatting. This transformation enhances the document's usability for hazard identification and risk assessment processes in industrial settings.

False Positive Management

Even high-performing models generate false positives (non-issues flagged as hazards) and, more critically, false negatives (missed hazards). For process safety contexts where missed hazards can lead to catastrophic outcomes-fire, explosion, toxic release-false negative management is paramount. Alert fatigue from excessive false positives can also erode operator trust and reduce the system’s practical value.

Solution: Implement confidence scoring systems derived from token log-probabilities and consistency voting across multiple model runs. Route low-confidence outputs to human validation workflows, focusing expert attention where it adds the most value. The chemical accident report extraction study demonstrated that confidence calibration flagged approximately 87% of errors via low-confidence outputs, making human-in-the-loop verification practical rather than burdensome. Human-in-the-loop verification is necessary to validate automated hazard identification outputs, and establishing clear escalation protocols-particularly for high-severity potential hazards-ensures that critical findings receive expert evaluation regardless of model confidence.

Regulatory Compliance Verification

ISO 45001 requires that hazard identification consider human factors, non routine activities, emergency situations, changes, and all affected persons. Singapore’s WSH Act imposes additional obligations for Major Hazard Installations. Auditors conducting ISO 45001 certification assessments or WSH inspections may be unfamiliar with LLM-based outputs and may require demonstration, validation documentation, and audit trails before accepting automated hazard registers as compliant inputs.

Solution: Build compliance verification modules directly into the system that cross-reference extracted hazards against the specific requirements of ISO 45001 Clause 6.1.2, Singapore’s WSHA provisions, and process safety codes. Ensure the system produces audit-trail-qualified outputs with full traceability-source document, extraction method, confidence score, validation status, and reviewer identity. HAZOP analysis examines each process node for deviations, and automated systems should replicate this systematic coverage to satisfy auditors that no operability study requirements have been missed. Engage regulatory bodies early in the implementation process to build acceptance and establish precedent for LLM-assisted compliance workflows.

Proprietary and confidentiality concerns also demand attention. Company documents-P&IDs, internal incident investigations, detailed design specifications-contain sensitive intellectual property. If using cloud-based LLM providers, data governance, encryption, data residency within Singapore, and non-disclosure protections must be established. On-premises or private cloud deployment may be necessary for the most sensitive materials, particularly given the strategic sensitivity of Jurong Island operations.

Conclusion and Next Steps

Automating ISO 45001 hazard identification in Jurong Island refineries through LLM-powered document analysis delivers measurable improvements in coverage, speed, and consistency-addressing the fundamental limitations of manual review while strengthening regulatory compliance. With demonstrated model accuracies exceeding 94% for consequence prediction and the ability to process thousands of documents in hours rather than weeks, these systems represent a practical advancement for process safety management in Singapore’s petrochemical industry. The technology does not eliminate the need for expert judgement; rather, it amplifies expert effectiveness by handling the volume-intensive extraction work and directing specialist attention to the findings that matter most.

To begin implementation:

  1. Conduct a document readiness assessment across your facility-inventory all safety-critical documents, evaluate formats and quality, and identify preprocessing requirements

  2. Select pilot process units with mature hazard registers that provide reliable baselines for validation against automated outputs

  3. Establish success metrics aligned with your current risk assessment process: target entity extraction F1 above 0.77, severity classification accuracy above 86%, and time reduction of 70% or greater

  4. Engage certified safety professionals and regulatory stakeholders early to build validation frameworks and audit acceptance

Related topics worth exploring include integration with digital twin technology for real-time hazard visualization, predictive maintenance systems that combine sensor data with document-based hazard intelligence, and multimodal AI models capable of analysing both text and visual data from P&IDs and equipment photographs. The development of benchmarks specifically for HSE compliance-such as HSE-Bench-signals that the industry is moving toward standardized evaluation of AI-driven safety tools, and early adopters on Jurong Island are positioned to lead this transition.

The image depicts a modern industrial control room featuring large digital display screens that showcase real-time hazard monitoring dashboards, including colour-coded risk maps and automated safety analytics for effective hazard identification and risk assessment. This advanced setup emphasizes the importance of process safety and risk management in monitoring and evaluating potential hazards in industrial operations.

Additional Resources

  • ISO 45001:2018 implementation guidance for refineries: Clause 6.1.2 hazard identification checklist covering all required consideration categories-routine and non routine work activities, human factors, emergency situations, design modifications, and contractor exposure

  • LLM evaluation frameworks for industrial safety: Key metrics include entity extraction F1, severity classification accuracy, causal chain completeness, confidence calibration, and false negative rates for high-consequence hazard types

  • Singapore WSH regulatory compliance: WSH Act obligations for Major Hazard Installations, SS 506 Part III process safety requirements, and risk management regulations for Jurong Island facilities

  • MOSAIC Ecoconstruction Solutions consultation: For guidance on ISO 45001 certification consulting and implementing safety management systems aligned with Singapore’s regulatory framework

Tags

What do you think?

Leave a Reply

Your email address will not be published. Required fields are marked *