Australia

As the development and implementation of Artificial Intelligence (AI) tools globally continues to rise, we are beginning to see trends in published research considering its benefits and drawbacks. Several bodies have developed rules underpinning its use in their environments, while others (including governments) are continuing to work on building further regulations. The author has reflected on the latest trends in this space, especially the use of AI by project management professionals. In the latter part of January 2024, the author embarked on an extensive review of AI's current abilities in delivering Critical Path Analysis (CPA) outputs by testing ChatGPT (v3.5), and in cooperation with others, on ChatGPT (v4), as well as several other AI project management AI tools. This article will highlight the limitations in ChatGPT's and other AI tools functionality in compiling CPA outputs. Using an example from Project Management Institute (PMI), the author will present these errors as well as a critique of the outputs.
The aim and objectives of this article are, through a case study, to conduct a critique on AI tool outputs on the specific scheduling technique of Critical Path Analysis (CPA). As part of wider research interest in AI and its use in project management and project control the author was reviewing a series of slides issued by the Project Management Institute (PMI) with the title 'Generative AI Overview for Project Managers – Resources' (PMI, 2023). These demonstrate the use of ChatGPT in project management, discussing the use of GenAI and generally DOs and DON'Ts in the use of GenAI.
In slide 12 of the PMI presentation a prompt is given (see Table 1 below) for ChatGPT to provide a Critical Path Analysis output for a hypothetical simple schedule of five (5) activities. The author will use the ChatGPT output provided in the PMI slides (see Figure 2) to highlight an interesting 'error' in the resulting outputs.
I am a project manager for a construction development project. The project has five activities:
Show me, in a table format, the dependencies between the tasks and their corresponding early start (ES), early finish (EF), late start (LS), late finish (LF) times; float for each task; and all the paths with duration. Solve this and show your work concisely with the critical path and all other paths in days. Highlight the critical path with the shortest duration and least float value of 0 in bold.
The ChatGPT output, as presented by PMI, is shown below in two parts Figure 1 and Figure 2 for ease of reading.
The author reviewed the output presented, and several basic errors were immediately obvious. Figure 3 presents the points in question, which are discussed in more detail below.
Some of the questions raised by the AI tool output are shown as discussion points 1 to 5:
The author considered that there must be an error in the PMI slides presented and decided to test the same prompt in ChatGPT (v3.5) during the latter part of January 2024.
The test was conducted a number of times repeatedly and despite using the same prompt, ChatGPT's (3.5) response changed each time and each response continued to present inaccuracies / wrong results.
The author considered that the results needed to be cross checked and two other individuals were asked to conduct the same test (in the same period – latter part of January 2024). Using the same prompt and in addition to testing ChatGPT (v3.5) they were free to check the outputs from any other AI tool they had available in order to achieve a wider perspective on the AI tool performance.
The two individuals in addition to ChatGPT (v3.5) used three different AI tools:
The results from all tests are included below, and the author has introduced some markers to highlight the resulting errors.
The results from Test 1 – the Project Management GPT's provided the most accurate output from the tests carried out; however, the graphic shows the following errors were identified:
Further discussion of these results will be presented later.
The author has highlighted a number of errors numbered above as discussion points 1, 2 & 3:
The author has highlighted several areas and considerations below as discussion points 1 to 5:
The author has identified the following errors in Figure 7.1:
Despite the errors discussed above and as can be seen in Figure 7.2, the tool attempts to validate its calculations by providing relevant sources.
Figure 7.2 Output from Test 4 – the PMI ‘aiassistant’ tool indicating sources used for the result.
The author has checked reference 2 – 'Project scheduling: improved approach to incorporate uncertainty using Bayesian networks' authored by Khodakarami, et al. (2007) – and this article has no relevance to the output the tool has provided. Was this a hallucination by the AI tool?
In addition to the above errors, there is also another consideration.
Practitioners and project management professionals know that the software tools that perform CPA use a calendar and, in the majority of cases, these are set up to have a working week of five days and a weekend. Therefore, it is known that activity durations should represent working days, and the scheduling calendar takes into consideration the weekend days off (Saturday and Sunday or Friday and Saturday). However, based on the outputs displayed above from both ChatGPT (Figures 4, 5 and 7.1) and the 'Bot' (Figure 6) do not recognise this.
Whatever happens in the Large Language Model (LLM) Machine Learning (ML) model(s) used, there is something fundamentally wrong with the processing of prompts and CPA calculations produce the wrong outputs.
Unfortunately, at this point in time, this is not a convincing case to encourage professionals to use the AI technology promoted in whatever form. Given the number of errors across the board, it is not just a case of asking professionals to use the technology with caution but ensuring that it delivers the required output(s) correctly.
In the past, similar mistakes have occurred with new technology being promoted as the answer to various issues too early. The author can see history repeating itself and practitioners ignoring and / or rejecting the latest technologies or perhaps worse, using the technology and obtaining the wrong results. A question here on this point; with the current technology what would have happened if the sample schedule was for 30 activities?
To provide a more accurate response to the prompts as a comparison, the author modelled the example in the scheduling tool MSProject, and the output can be seen in Figure 8 below in a 'straightforward' logically linked barchart view.
The use of a normal 5-day working week calendar as well as the normal Finish-to-Start logic links, as per the example, provides a view of the appropriate ES, EF, LS, and LF dates. The column 'Total Float (Slack)' shows the 7-day TF against Activity 3 (as is also shown in the Project Management GPT output, however, without errors in the start and finish dates).
The results presented from a number of AI tools indicate a clear misunderstanding of CPA theory, including basic calculations and important outputs such as TF, ES, EF, LS, and LF.
If this happens with a relatively simple case, how could we trust AI's outputs to the CPA of more involved schedules with more activities and more intricate logic links, as in real life. In reality, we also introduce leads and lags, different calendars for different resources at different locations, etc., and the process of scheduling and producing a CPA output is much more intricate.
The results presented above do not allow for any confidence in the AI tool output(s), regardless of the tool used.
Another point to consider is the process between understanding the requirement, raising a query in any AI tool, and using the output. According to the PMI there should always be a 'Human in the Loop'. Using AI requires a complex iterative process where prompts must be set out clearly and accurately and outputs thoroughly reviewed and tested before they are released to the audience.
Why are we and should we be using LLM to carry out intricate duties, such as CPA?
Over the last few decades, we have developed appropriate software tools to deliver CPA outputs for projects of any type. What is the use case for re-inventing the wheel and possibly using the 'wrong' AI model or using AI in inappropriate ways to reproduce these outputs? How can an LLM model understand schedule network outputs that cannot, at the moment, be calculated and assessed?
Shouldn't we be introducing the appropriate ML methodology within the existing scheduling software and therefore enable the AI tool to 'learn' from other similar schedules? Even better, what if we enable the ML tool to access previous real-life outputs, from real project schedules, conduct analysis of the relevant schedules and then 'instruct' it to provide us with possible outputs/scenarios?
Therefore, the author considers as more appropriate that AI tools / technology should be developed in a way to enable them to work from within the scheduling software tool(s). ML models (LLM in this case) that exist for generic purposes should not be used to provide outputs to specialist questions.
Similarly, AI implementations can be integrated with other project management software tools, for example, estimating, cost management, and contract management, to enable scenario building as well as improved performance and output from the relevant processes, rather than allowing them to bypass them.
We are currently in the 4th Industrial Revolution (4IR) era, and changes, as predicted by all bodies, are rapid and technically challenging. Contributing to the challenge is the scale at which various AI LLM tools have been rolled out within in the last two years and our ability to grapple with this. Despite AI having been a part of our lives for decades, the seeming 'uncontrollability' of it has caused mixed reactions in terms of the effects in our professional as well as personal life.
The issue that we are faced with is how we can best, and most of all ethically, utilise the power of the tool(s) we have developed in order to deliver an improved output. By cutting corners and bypassing processes we have developed we will not achieve a coherent output. We need to muster our collective expertise and build tools that integrate AI using existing capabilities.
The power of AI can be harnessed by applying it to support the relevant tools and not by superimposing it. Asking LLM models to deliver a Critical Path Analysis is neither appropriate nor correct. However, integrating ML tools, and perhaps a type of LLM model(s), within the scheduling software tool, or any other relevant project management software tool, and then asking it to deliver possible scenarios, is the correct approach to surf the 4th IR wave. We will still have to have the 'Human in the Loop' to interrogate the output and advice, and therefore our next steps must be to build, educate and inform skilled personnel on how to achieve this.