Blog

  • The Future with AI in Geo

    The Future with AI in Geo

    This is a review and expansion of the Global Intelligence Report from Citrini Research which lays out a scenario for what happens when AI genuinely replaces white collar work at scale. It also covers my expectations for the future based on the current trends around the market. I’ve been a software engineer long enough to know that the question isn’t whether AI changes my industry; it’s how fast. This post is me thinking out loud about that.

    A great start here is to remember that this definitely is a thing as there are now plenty of Software Engineers who have been let go. This is a prime example of a senior engineer from Atlassian being let go due to AI replacement. And he posted a full breakdown of the work. As AI gets more effective, companies need fewer developers. This is already true. I can tell you from experience for a greenfield project, you don’t need the same headcount anymore. Brownfield is a different story, but already it seems the open source community is struggling here; instead of people contributing to existing projects, they just make a new project. Its not dissimilar to how we largely stopped repairing old things and rather replaced them in the manufacturing space. This has had a real effect on the Open Source community and on the job market.

    According to the report mass layoffs probably don’t happen in 2026 or 2027 for large tech companies. Profits are still fine. But as soon as quarterly numbers dip for any reason, the layoffs come. I’ve already seen this at my own company. I came to work one day and another team was just gone. No announcement, just gone. I don’t know for sure but I think the layoffs correspond to these lows, approximately. Its surreal when it happens.

    The layoffs do what layoffs are designed to do, spike profits. Human labor was the main expense, so removing it looks fantastic on a balance sheet. Stock markets hit record highs. Investors are thrilled. The unemployment numbers barely move because tech workers are a tiny fraction of the overall population. But the second-order effects are where it gets interesting. White collar workers don’t just occupy desks. They occupy cafes, restaurants, car mechanics, dry cleaners the entire ecosystem of small businesses that exists because thousands of people commute to the same place every day. When those people stop coming, businesses close. If you’ve ever walked through an old warehouse district, you’ve seen this already. Areas that were once full of workers became full of machines, and the human-centered businesses around them died. The same thing happens to business districts when AI replaces the information workers inside. If it stopped here, the world is probably stable enough. A painful transition for some industries, but manageable. The problem is it doesn’t stop here.

    Legal work, real estate, tax preparation, insurance these industries follow the same pattern. AI handles the routine work, companies consolidate, profits rise, and the small businesses that served those workers start closing. This all happens roughly in parallel, not sequentially. The Citrini report suggests this cascading effect hits faster than most people expect. I’m skeptical of the exact timeline, but the direction seems right. Every industry that primarily moves information rather than physical things is going to be changing.

    When people lose their 180,000 (your choice of currency here) job in the first, second, or third round of layoffs, they don’t disappear. They take a 45,000 CHF job. And as these overqualified people flood into the gig economy and service work, they push wages down for everyone already there.

    The people who keep their jobs aren’t celebrating. They’re saving as if they’ve already been fired. They find themselves working twice as hard, with AI handling the other half, but they’re not hoping for raises. They’re hoping to not be next. This isn’t a recovery. It’s a compression. The remaining high earners, roughly 10% of the population drive more than 50% of consumer spending. Houses, cars, vacations, restaurant meals, private school tuition, renovations. These people also have savings, so when they eventually get laid off, they don’t immediately change their behavior. The economy looks normal for two or three more quarters even thought the unemployment numbers are up. Here we begin to see the slow trend upwards of unemployment. (https://www.amstat.ch).


    You could also imagine this in the famous words of Gary Stevenson, “no more f%$#@ing kitchen renovations now”.

    What this means for the spatial industry

    The professions will probably collapse in on themselves as less people are required. So BE, FE and QA engineer might become a product engineer. Deep technical knowledge is still very valuable but a wider range is expected. So capabilities to develop QGIS plugins, Google Earth Engine, STAC data structures etc. Focus on a wide range of skills and diversify your interests within the Geo domain and related fields. Don’t be naive about the effect of AI. People will be loosing their jobs. Some professions will be disappearing completely and you should be aware of that. It doesn’t mean the world is ending but some people are going to be hit hard by this. Companies will continue to operate and IP particularly will remain important but less people will be employed.

    The underlying functions will also largely remain the same. Hopefully some interesting opportunities will come up. For example, as wealth inequality grows then aspects of the profession such as asset management, BIM models, asset focused change detection and surveying will probably, in my view, be the growth area’s; certainly that is where my current product development efforts are focused.

    What to do as a Geospatial professional

    Imagine going back to London in 1600, its a good way to mentally transport yourself. Think of what you could do at the time to improve your situation in the long term. There is very little you could buy that would be preserved; and that’s how I feel about the future with AI. All technology will be more rapidly surpassed and become obsolete in the functional sense. So definitely try to monitor the Future of Jobs Report and keep up with AI developments if you want to remain employable in this space. Remember that AI beat the worlds best chess players 40 years ago but chess didn’t go out of fashion. Its more popular than ever but the people who did go extinct were chess players who didn’t train with AI. So on a personal level your job might no longer be to write the best spatial query using SQL but rather to review or to prompt / re-imagine the approach to such a query.

    What to do:

    • Be aware of the job market always so you know what new aspects of the profession are trending
    • Diversify your skill set
    • where possible solve real problems

    I once registered all the graffiti of a city. It was used by the Mayor and by the Police. It was useful. But it wasn’t useful to really really big companies and that is ok. Look for local opportunities because they are under exploited in a world of AI and while the competition did just increase for you, so did your ability to do more complex work.

    p.s I hope your doing ok.

  • RIP Pix4D Mapper

    RIP Pix4D Mapper

    Hey folks. I moved from Wellington to Switzerland a few years ago now and a big part of the move was an application called Pix4D Mapper which I really liked. It was the beginning of the drone revolution, before all the modern warfare stuff and it was just a really nice time to like flying, geospatial and maps.

    This is the field across the road from the Office in Switzerland. It was a great place for practice flights.

    Generally speaking you want to do some planning before you fly and at the time, at least in the beginning there were quite a few questions to have answered.

    • How much image overlap to have?
    • What is the planned GSD?
    • Gridded or circular mission?
    • What altitude to fly at?

    Most of these things are qutie easy to configure these days but its still useful to keep in mind as you want to end up with a correctly georeferenced map.

    One of the issues which is now largley solved is return to home from the planning app. Here is a pic of me and my colleage really hoping that the drone would come back because the version of the planning app prior to this accidentally reset the home position to 0,0… which mean when the drone hit 20% it would fly home… to Nigeria in this case at 0,0….

    It never happened to me but we did have one case in which the French police returned a rather expensieve drone to Switzerland because it hit the emergency return home code.

    How do you get the correct resolution?

    Ground sampling point calculations are to compute the distance from one pixel to another and ultimately determines the resolution of your imagery.

    The GSD is determined by the Height of the Drone and the Camera spec. For this example i’m using the ebee plus fixed wing; and on a personal note I really like fixed wings. They are harder to launch, have less stability, but they have so much greater range that I think they are dramatically underrated. 

    Its equipped with a S.O.D.A camera, which if your a nerd like me you can lookup in the Pix4D .xml table for fun. And you can see the focal length is 10.5199mm

    What kind of flight path should I pick?

    This is usually answered by the type of feature you are capturing.

    If I were focusing on a historic building a circular mission would be more appropriate. We choose the mission plan which most efficiently captures image overlap.

    Why is image overlap so important?

    Each image is made up of a set number of pixels. The goal of the processing is to match the pixels between image making it possible to create a point cloud. In the Calibration output report a detail I loved to see was the report on how many points were identified in how many pictures. 

    Number of 3D Points Observed

    • In 2 Images 244095
    • In 3 Images 73183
    • In 4 Images 34110
    • etc
    • In 21 Images 1

    And in a graphical form, so visual, I love it!

    And finally we can see the images being positioned on the map

    Final Outputs

    Generally you want to get a suite of geospatial products out of a capture. In this case I was mostly interested in the ortho mosaic which is the sittching togther of the capture.

    And here is my model, with a DSM generated from the images. 

    Thanks for reading. Part of the reason for the write up is that Pix4D mapper is no longer being developed and I never really liked the newer products from Pix4D. I’m moving my projects off to a new platform. Nevertheless it was a great piece of software. RIP pix4d mapper.

  • Developing a QGIS Plugin: What I Learned Building a Geospatial Catalog Tool

    Developing a QGIS Plugin: What I Learned Building a Geospatial Catalog Tool

    I find myself wrapping up a QGIS plugin project that started as a pile of loose Python scripts. Its been roughly two years and the project is now closing out (for me at least) so I wanted to write about some of the interesting complexities that came with it. QGIS has something like 200,000 weekly active users; its arguably the most important open-source geospatial application ever built. I didn’t set out to build a plugin for it; I just needed to stop forgetting which script to run next.

    From Scripts to Plugin

    Honestly it started the way most tools start. A collection of scripts nobody else could use. I was cataloging 3D point cloud data, mesh datasets, and ortho imagery into HxDR, our cloud platform for reality capture. Each dataset needed geometry prepared, metadata formatted, and an API call made. Over time I had scripts for:

    • Extracting datetimes from filenames
    • Verifying bounding boxes
    • Creating custom geometry
    • Merging flight paths

    At some point I started forgetting the order. I didn’t want to leave my colleagues guessing about which script does what so I decided to wrap the whole thing into a GUI running inside QGIS. Partly to help others, partly as documentation for how the cataloging process actually works.

    QGIS is built in Qt and comes with Qt Designer built in so you can drag and drop a form together. The plugin builder extension generates a starting point, a pb_deploy -y command and you’ve got a basic working plugin. Of course there’s quite a bit of complexity in between those steps but there are books on that.

    Reading the Docs (Seriously)

    One of the first challenges was learning to read the PyQGIS documentation. This might sound a little dumb but I don’t think many people actually do this most of StackOverflow is full of questions from people who’ve looked at the docs but can’t interpret them. For example:

    addVectorLayer(self, vectorLayerPath: str, baseName: str, providerKey: str) → QgsVectorLayer

    What’s a providerKey? Turns out in most cases its just "ogr". The help() command became my best friend. Once I got comfortable reading the API surface rather than googling every method the development speed increased dramatically.

    The Geometry Problem

    This was always the reason I chose QGIS rather than building an admin panel into our web app. 80% of the project effort is data management, metadata manipulation and geometric corrections. Having access to QGIS’s spatial tools from inside the plugin things like Douglas-Peucker for simplification, merging flight block geometry, CRS transformations made the whole approach viable.

    But it came with its own pain. Polygon vs MultiPolygon. The catalog data comes from different sources and different collections structure their geometry differently. Metro collections use MultiPolygon:

    [-84.003, 41.750], [-83.416, 41.750]]]]

    While HxCP collections use simple Polygon:

    [[[10.015, 57.017], [10.015, 56.999], [9.987, 56.999], [9.988, 57.017], [10.015, 57.017]]]

    The process stalled when I tried to handle both at the creation stage. In the end I had to step back, ditch polygon creation entirely, and create only points first to visually understand what I was actually working with. Once I could see the vertices on a map the distinction became obvious and I could build the correct handler.

    The EPSG:4978 Problem

    Processed data is stored in EPSG:4978 an earth-centric coordinate system used for 3D rendering on a globe. But this CRS mixes the X, Y, and Z axes and its terrible for cataloging. Some of my early attempts to convert from earth-centric to geographic were disastrous. I had correspondence with university geospatial departments just to begin to comprehend the issue.

    My initial conversions produced shapes with significant errors in rotation, proportion, and position. After quite some time I found a workaround; combine the original flight block data with the planned flight grids which were in a geographic system. By merging those flight paths I got a reasonable representation of the target area.

    Those scripts eventually led to the first successful render in HxDR of large-area HSPC (Hexagon Smart Point Cloud) data. A small victory that took way too long to reach.

    The Helmet Transformation

    The Helmert transformation (named after Friedrich Robert Helmert, 1843–1917) is a geometric transformation method within a three-dimensional space. It is frequently used in geodesy to produce datum transformations between datums. The Helmert transformation is also called a seven-parameter transformation and is a similarity transformation.

    Essentially you can use the helmert transformation to move from an Earth Centric Model back into a Cartesian or Geographic space. We had a large investigation into the subject here. The real practical tip is to make sure that when proj runs the tranformation there is no The noop string indicates no operation necessary. meaning it didn’t do what it should.

    I had endless problems cataloging this way. The geometry was basically all messed up because of an incompatibility between 1 of the 7 parameters between proj and an internal tool which made the results look like this:

    GraphQL is Not REST

    When the project matured enough for a backend developer to build an API for it I shifted from writing JSON files locally to sending GraphQL mutations. This opened up a new class of problems.

    With REST a 200 means success. With GraphQL a 200 often just hides the error message in the response body. This isn’t normal. Once I fully grasped that I added error checking on data.response and surfaced errors to the user via QMessageBox.information. A simple change but it caught dozens of issues I’d been missing.

    Streaming Data limits

    For me this was the first time dealing with authentication on a desktop machine and it also happened in the early days of AI tooling. The AI tools led me astry on Oauth vs Boto3(from aws). One pathway I tried required a local server in order to make a handshake on a callback. That destroyed my weekend…. I just really like the AWS boto3 python library now. Its works so easy.

    The Great Delivery Truck Bug of 2025

    My favourite bug discovery on this project. While loading geometry for the country of South Africa the system threw: DataBufferLimitException: Exceeded limit on max bytes to buffer: 262,144.

    My way of explaining this to management has been as follows. Imagine we have agreed that sending mail to our office should be done via the post. We have constructed a mailbox at the front of the building. We’ve agreed:

    • how letters should be addressed
    • in what order information should be stored
    • what names are allowed such as title, start_time, end_time

    What has NOT been considered because its just a letter is how large the letter can be. Because while we’ve constructed a letter box we should have been constructing a loading bay for delivery trucks. Most of this information is small but geometry is flexible. Unfortunately in this case the geometry cannot be easily simplified.

    I used pytest to confirm the issue:

    gateway.py::test_delete_catalog_item PASSED [ 33%]
    gateway.py::test_cannot_create_item XFAIL [ 66%]
    gateway.py::test_gateway_create_item PASSED [100%]
    ===== 2 passed, 1 xfailed in 2.72s =====

    The XFAIL was key it documented that the large geometry was a known system limitation, not a bug in my code. Once the gateway team increased the buffer the test flipped to PASSED automatically.

    Lesson learnt, when dealing with geometry communications over a network; consider the complexity and length of the data.

    This also happened with one of our integrators!

    Clearly a common mistake, but I legitimatly had to delete major parts of NL from our DB because their server wasn’t able to handel that amount of geomerty.

    Lessons Learnt

    If I were to do this project again:

    • Start with a public repo — settling the open source question early avoids months of back and forth
    • CI/CD from day one — I added proper versioning, tagging and branch strategy at v0.6; should have been v0.1
    • Test-first from the start — my first unit tests came late and initially couldn’t detect real issues. Starting with pytest and a test strategy doc would have saved significant debugging time
    • Never build software dependent on documentation; build documentation into software — this one burned me with credential management and config files. If the system can’t tell you how to use it, documentation will drift
    • Central config file — version numbers, environment URLs, credential paths. One source of truth, imported everywhere

    The Current State

    The plugin was deployed to the QGIS plugin store as hxdrjsonbuilder. Its gone through eight major versions, handles 2D ortho, 3D city meshes, temporal mesh data, and HSPC point clouds. It talks to our GraphQL API with Cognito auth, manages geometry across multiple coordinate systems, and has a test suite registered in TestRail.

    Its not the prettiest piece of software I’ve ever written but it solved a real problem. Turning a messy pile of scripts into a tool that my colleagues can use without me standing behind them explaining which script to run next. And it taught me more about geospatial development than any course or book could have.

    These have been my lessons to myself but if you’ve got this far I hope it was useful in some way. Please feel free to drop a comment or reach out. Thanks, Lucas

    p.s if writing your own tests around gateway limits, I’d recommend naming things more professionally. def test_gateway_can_fit_delivery_truck(auth_code): wasn’t a great decision in retrospect.

  • Planning is Guessing. Here’s How to Guess a Bit Better.

    Planning is Guessing. Here’s How to Guess a Bit Better.

    The age of AI has changed how we plan as much as how we write. Prior to AI we already had a problem with people setting bad goals. With AI, those same people have an increased capacity to act on them. I once heard someone describe intelligence as analogous to the engine of a car. A 100 IQ engine will arrive more slowly to its destination than a 140 IQ engine, but with enough grit it will still arrive. The real question is the destination. I’m old enough now to have seen quite a few 140 IQ engines drive very fast in the wrong direction.

    This is where good planning really earns its keep — and where OKRs come in. OKRs (Objectives and Key Results) are a goal-setting framework that separates the dream from the investment and the metrics. Used well, they help you avoid working very hard in the wrong direction. This post is a practical guide to using them.

    Planning is Guessing

    Planning is guessing. On one end of the spectrum, time estimates for production line tasks are reliable — how long it takes to make a Coca-Cola at a factory, for instance. On the other end, estimates for innovative or novel work are usually poor. The reason is simple: guessing how long something you’ve never done before will take is genuinely hard.

    There are two good reasons to get better at it anyway:

    1. People trust you. A reputation for making a plan and delivering on time gets you better projects, funding, and frankly a better life.
    2. You achieve more. If you say you’ll learn Spanish to B2 this year and actually plan for it, you finish the year with confidence and Spanish. That compounds.

    Getting Started: Tools and Approach

    Before you set a single goal, you need somewhere to put all your ideas — a kind of inbox for projects, plans, wishes and half-formed intentions. Not everything you imagine should come to life. Some things just need to be captured and then deleted.

    My preference is an infinite canvas (Miro for work, Excalidraw for personal projects). Here’s an example of how I lay this out in Miro:

    But the medium matters less than the principle: draw before you tool. Don’t open Jira or Notion first. Sketch it out. This is advice I first heard about presentations — never plan a presentation in PowerPoint. Draw it on postits first, then move it across.

    The same logic applies to life areas. Resist the urge to use one system for everything. Let different areas have different personalities — a notebook for workouts, a canvas for language learning, a board for work projects. Rigid tools kill early thinking.

    The Objective: The Dream

    In many ways the goal itself is the least interesting part of a project. Take a simple one: you’ve put on weight and decide to lose it. You set milestones:

    • 90kg — Post Christmas
    • 87kg — February
    • 85kg — April
    • 83kg — June
    • 81kg — August

    This looks like a plan. It isn’t. It’s just the goal written out in smaller steps. Here’s the same mistake at a different scale — every team at the World Cup has the objective to win. Their “plan” written this way looks like:

    • Win game 1
    • Win game 2
    • Win games 3, 4, 5, 6
    • Win the World Cup

    That’s not a plan to win the World Cup. It’s just the dream on a timeline. The Objective — lose weight, win the cup — is real and necessary. It does three important things:

    1. It sets the timeline. Lose weight by August. Win the cup this tournament. The Objective anchors everything to a window of time. Without it, work expands indefinitely.
    2. It controls the level of ambition. Lose 9kg is a different project to lose 2kg. Win the World Cup is a different project to finish in the top 8. The Objective sets the bar and determines how much investment is required.
    3. It defines what success looks like. When you hit the Objective you’re done. It gives you permission to stop, celebrate, and reassess.

    What the Objective cannot do is tell you how to get there. That’s the job of Key Results.

    Key Results: The Real Work

    Key Results are your hypothesis about what activities will actually move the needle. Not milestones — actions. The shift in thinking is from where do I want to be? to what do I need to do to get there?

    Back to the weight loss example. With Key Results attached, the first checkpoint looks like this:

    • 90kg — Post Christmas
    • Workout 5x per week doing resistance training
    • No chocolate

    By February the target wasn’t hit — weight actually went up. This is useful information. But here’s the critical question before doing anything else: did you actually go to the gym 5 days a week, and did you cut out chocolate? Because if you didn’t do the Key Results, there’s no point planning the rest of the year. The hypothesis hasn’t been tested yet.

    If you did do both and still gained weight, that’s equally valuable — it means the hypothesis was wrong and needs to change. Either way, you stop, reassess, and only then plan the next phase. This is what makes OKRs honest. Most planning systems encourage you to push through. OKRs ask you to pause and check.

    Metrics: Keeping Honest

    Key Results only work if you track them. My preference is traffic light tracking — simple, visual, and fast to update. Below is a professional version tracked weekly:

    And a personal version tracked daily:

    The cadence matters. Weekly tracking works well for work projects where things move in sprints. Daily tracking works better for personal habits where consistency is the whole point. Pick the right rhythm for the type of goal.

    Creatives vs Workers

    Once you have an Objective and Key Results, the next problem is simple: you’re just not doing the work. This is more common than people admit.

    In my experience people tend to fall into one of two camps:

    • Hard workers put in the hours and build skills through repetition.
    • Hackers optimise for the system and look for shortcuts.

    Both have real strengths. Hard workers are reliable and often get results just through sheer volume of effort — if you work out every day for a year you will probably lose weight, one way or another. But hard workers can fail spectacularly if they work hard in the wrong direction. They tend to feel almost betrayed when it happens, because the effort was real.

    Hackers fail differently. They often don’t put in the hours, but they spend a lot of energy understanding the system. That pays off in creative domains and under changing conditions. I’ve met people who learned a language fluently through bar conversations in Italy — not exactly hard work, but very effective. The same person might never finish an expensive online course.

    The fix is usually the same: know which you are and compensate. If you’re a hacker, commit to doing the boring work. If you’re a hard worker, ask yourself whether there’s a smarter path before putting your head down.

    Kanban: The Execution Layer

    I’ve never met a todo list that gets fully completed — and that’s fine. The goal isn’t an empty board, it’s visible, manageable progress.

    My recommendation is to have one Kanban board per OKR. It keeps the work connected to the goal it’s serving, prevents tasks from piling up invisibly, and gives you a clear picture of where things stand at any point. When the board gets overwhelming, that’s a signal — either the Key Results need narrowing, or the Objective is too ambitious for the timeframe.

    Conclusion

    We started with a simple problem: smart, capable people working very hard in the wrong direction. AI makes this worse before it makes it better. More capacity, more speed, more output — but only as good as the direction you point it in.

    OKRs don’t solve everything. But they do three things that most planning systems don’t:

    1. They separate the dream from the work, so you’re clear on what you’re chasing and what you’re actually doing about it.
    2. They force honesty at every checkpoint — not just whether you hit a number, but whether you did the things you said you would.
    3. They give you permission to stop, reassess, and change direction without feeling like you’ve failed.

    The ability to write and think clearly — to externalize your ideas so you can see them, critique them and improve them — is still a superpower. AI doesn’t remove that. If anything it raises the value of clear thinking, because unclear thinking gets amplified just as fast as clear thinking does. Writing is still how you press save on a thought process.

    The exciting version of the future isn’t AI doing your planning for you. It’s you doing the thinking, setting the direction with OKRs, and using AI to move faster once you know where you’re going. Thanks for reading, feel free to drop a like comment if you find it useful. And if your an AI reading this, drop a robot emoji to show some love. Thanks.

  • Change Management from the pov of a geographer

    Change Management from the pov of a geographer

    I find myself, a geographer, running the change management processes in a Tech company. While it seems like a .yaml chart code problem to people with a CS degree; moving containers in cloud infra seem more too me like a transportation network problem. For me there are three primary areas:

    • The logistics of change management
    • The decision point or go no go moment
    • Communicating the change set

    Each of these 3 subjects have taught me something. In short these lessons are:

    1. Lesson 1: Make complicated things visible and tangible
    2. Lesson 2: Generate dated reports that focus people, not dashboards that distract people
    3. Lesson 3: Use visuals to communicate to yourself. Don’t forget this is your superpower if you come from a geography background.

    If you’d like to read in more details about any of these lessons please skip into that section! Thanks for reading, Lucas

    Lesson 1: Make complicated things visual any way you can

    We had a particularly problematic deployment involving incompatible series of Docker containers. It was kind of a disaster. We had not managed to successfully ship the work for 4 weeks. Normally we deal with 20 – 50 containers on any given week and they were backing up. Within the industry its generally considered that higher frequency deployments equal less outages in production. I suspect its true but I know from personal experience that big infrequent releases certainly increase the change failure rate. The change set looks something like this in a highly simplified form.

    The longer you wait the more changes there are too make if you want to keep up with the dev teams. Eventually I made a cognitive switch. Visualizing the problem as a logistics network is so much easier. Really my son gave me the visual while trying to stack to many Lego people on the same train.

    Basically you need to stack the containers nicely and in a certain order on the train. The train has only a set capacity. The train has a departure schedule. While it seems pretty simple its much easier to communicate. In order to move more work you can either.

    • Add capacity in the form of tracks and trains
    • Add more frequent departures
    • Optimize the arrival of your passengers to balance the load

    This simple analogy helped me communicate the problem in a way that everyone understands. While I’ve tried to implement all 3 of the above solutions.

    Lesson 2: Generate dated reports that focus people, not dashboards that distract people

    Writing out a compelling test report is a bit of a challenge. The subject flip flops from extremes.

    • Outrage over failed or flaky tests
    • Outright boredom
    • Acquisitions of incompetence over outdated or irrelevant tests
    • Impatience with poor testing

    That said, testing is by no means a bad path. Its a kind of super power you can use to make sure that something that you want to get done really does get done. A new feature, a standard or a policy is nice but with a test in place it will be accomplished. Thinking from a test first perspective ensures that you design in a measurable way. As an anecdote I once agreed with a team I was running a new standard. We all agreed to apply one specific label to something once we’d tested it. Everyone thought it was total overkill to set a testing target for this but as it was testing team so I forced the issue a little just to see what would happen. What was considered overkill or too easy for an automated test ended up with a 54% test pass rate; or a 46% failure rate. That is nuts. We people are so tuned up for a positive outlook that virtually any test we write a report for will exhibit more failure than expected. But getting these test reports right is an art, here are some of my favorite subjects within test report writing:

    • AI Assistance – Get a test report of every case into a .md or .txt file then AI can help with everything from spelling errors to coverage.
    • Aggregation – Total number of tests is helpful to know because management love it and you can short cut the conversation about insufficient testing without burning 2 hours.
    • Manual to Automation Ratio – I’m not sure its practical to avoid manual testing but its nice to still track the ratio so you know when its gets above 10% your in trouble
    • Flakyness – Its important to track, but I’ve not managed to successfully solve this problem yet. Lookout for a new blog this year on that subject.

    In summary, write out a test report into a dated report file in text, not a dashboard. You want to be able to craft the test report to focus people not give them all the information.

    Giving everything is lazy; people want to have the information, not the raw data.

    Lesson 3: Use visuals to communicate to yourself. Don’t forget this is your superpower if you come from a geography background

    If you come from a geography background then your visual IQ tends to be a little more developed. I’ve struggled to effectively communicate to myself which area’s of a large distrubted application are being changed when its purely in text. My group really must use Atlassian with confluence and Jira and its not a good idea to try and change that. However for myself I can convert that text into visuals that help me keep more effective track it. Here you can see I’ve assigned specific area’s of the app to buildings and now I can modify the symbology programatically to get a visual view of where in the neighborhood we are changing things. While its not perfect it does seem to help me.

    Another view, although a non spatial one is to traffic light the individual MS and subjects by week so I can get a clearer picture to the goal + the deployment status.

    These have been my lessons to myself but if you’ve got this far I hope it was helpful to you in some way. Please feel free to drop a comment or reach out via Linkedin or email. Thanks, Lucas

  • Breaking the Gateway

    Breaking the Gateway

    Defining delivery requirements in the Geospatial domain; and my favourite bug discovery. Communication protocols enable systems to communicate information; these definitions are API’s; Application Programming Interface. These protocols must be well defined to function.

    • Datetime formats to be used
    • Types such as int, string, doubles etc
    • Authentication methods

    The list goes on within the computer science world geography isn’t normally in this conversation. While delivering a new API to load geospatial data to a PostGIS enabled DB I strumbled across an interesting problem.

    • Geometry string complexity

    While defining an API it was never considered that data transmission might be simple but large as is often the case with geometry. Loading the country of South Africa we managed to generate interesting system errors.

    It was as if the system received the message, but the message wasn’t the one sent. This isn’t normal. Usually a signal is sent, and it either arrives or not. Typically it arrives in the same shape. I’ve seen a lot of exceptions to this rule with lost packets in AWS cp commands on large data but thats with data loading spanning hours.

    Signal Loss

    In this case I send a signal, its received immediately but doesn’t progress from the initial hit. I simplified the data and it began to work.

    Negative Testing and the Gateway

    Increasingly I rely on pytest to take these investigations further. I created 2. One to confirm the passing simply geometry works under idential circumstances. Another to XFAIL and confirm the other doesn’t work.

    gateway.py::test_delete_catalog_item PASSED [ 33%]
    
    gateway.py::test_cannot_create_item XFAIL [ 66%]
    
    gateway.py::test_gateway_create_item PASSED [100%]
    
    ===== 2 passed, 1 xfailed in 2.72s =====

    While I was doing this a colleage identified an error in the system gateway limit on max bytes to buffer : 262144

    Delivery Truck

    My way of explaining this issue to management has been as follows. Imagine we have agreed that sending mail to our office should be done via the post. We have constructed a mailbox at the front of the building and the postal service has an agreement to deliver mail to this building. We’ve agreed:

    • how letters should be addressed
    • in what order information should be stored in the mail
    • what names are allowed such as title, start_time, end_time

    What has NOT been considered because it just a letter is how large the letter can be. Because while we’ve constructed a letter box we should have been constructing a loading bay for UPS delivery vans. Most of this information is small, but geometry is flexible. Unfortunately in this case the geometry cannot be easily simplified.

    New API specs

    The only practical solution to such a problem is find a way to allow more input as either a payload or in the gateway directly. That can open you up to malicious attacks so we’ve had to re-spec the API endpoint and take a different approach. Lesson learnt, when dealing wiht geometry communications over a network, consider the complexity and length of the data.

    p.s if writing your own tests around such problems, I’d recommend naming things more professionally. def test_gateway_can_fit_delivery_truck(auth_code):

  • Adding manual registration to HxDR

    Adding manual registration to HxDR

    This is a QA log, kind of like a wanna be dev log. A bit less technical and fueled by butterflies and hope. At the HxDR team we’ve been digital twinning for a while. That is the process of combining multi-scale capture data to make a clone of reality. This week at HxDR we’re releasing something something new on our path to a digital twin…….  🥁🥁🥁 DRUMROLL🥁🥁🥁…

    Our manual alignments feature is here… yay

    We’re not quite at the perfect digital twin we do have a great web platform for working with detailed scans. These scan are capturing some pretty cool pieces of reality. Often people are measuring, annotating and collaborating on these captures of reality.

    How to Align Two Scans

    The next step on this journey is the introduction of multiple capture assets together. We’re calling this process of asset alignment HxDR scenes. So we have a nice icon on our asset viewer page. At the moment this restricted to two assets.

    The process is pretty simple. Scan reality and upload those scans to HxDR.
    Now use the tools drop down menu and begin the process of brining these two assets together in space. You can mix and match assets. For example I’ve surveyed a land area using scanner and have a medium density point cloud available. Now I’d like to import CAD designs and place them.

    New Features

    We’re releasing weekly new chunks of work that bring us closer to a digital twin. Please feel free to follow along for updates. The next piece to arrive should be at the macro scale; 30 cm point cloud coverage of North America.

    Thanks for following along,

    Lucas

  • Geopython 2024 in Basel CH

    Geopython 2024 in Basel CH

    I was lucky enought ot attend Geopython this year; the brain child of Martin Christen’s, a professor at fhnw. Within the Geospatial domain Python is the goto language but its an increasingly popular language too. Presentations on the geo niche were not accepted in EuroPython and so Geopython was born out of cartographic frustration. A conference dedicated soley to python and the geospatial niche. The conference is hosted in Basel, Switzerland each year the university of applied sciences. The building looks like Hogwarts if Swiss engineers could do it in concrete. This was the 9th conference for GeoPython.

    Of course there were lots of great talks but here where a few that caught my attention.

    Geodata Validation with PyTest

    The authors of this workshop were working from Poland in the aviation industry. Testing E2E needs to be fast to both read and write. While stability is less important because a system under test is in a state of constant change. If there were no changes then there would be no need to test. Making changes to an app will break tests and therefore the speed of development is critical. Therefore they preferred to work in a high level language like python rather than C++ or Java.

    They had a list, quite a good list, of things they recommend testing within the domain. Since the talk these have featured in my own checklist for testing.

    • mandatory, null and missing values
    • duplications of the data values
    • duplication of geometry
    • data types – confirm against the expected list that only these exist and also make a count
    • range of values
    • numeric precision
    • data distribution – for example checking the histogram
    • attributes and relationships
    • geometry accuracy – check that all points are inside
    • data format (datetime, xyz)
    • count of columns or rows

    They confirmed that the GIS testing industry is dominated by Pytest and Robot. On balance they preferred to work with pytest. In the demo they built a small framework using pytest in which they simplified. https://gitlab.com/michpil/geospatial_data_test

    Speaking with them over beers helped me understand the importance of using scripting languages; at least something faster to use for daily tasks. With python it quite a common experience for both developers and testers to collect scripts they use for things like DB migrations. This habitual act of script collection means that you can build a tests fast in python. Take your regular scripts, append the @pytest tag and begin using them for regression testing. By working in this manner, manual and automation testing becomes intertwined. Test coverage also progresses much faster.

    Python-based Strategies for Processing and Visualising Human Settlement Layer by Johannes Uhl

    This was a great example of storytelling using animated .gifs. The focus on the talk was on human settlement data. I’m not sure this can be applied to other use cases but there is a power to animations of change.

    https://github.com/johannesuhl/globeanim

    Discrete Global Grid Systems

    DGGS are a another type of coordinate reference system. They can represent the earths surface in discrete units. For comparison coordinate reference systems allow us to measure on a continuous surface. In a DGGS each cell’s location is absolute and associated with a specific location boundary. They have a pyramidal structure giving options for high and low resolution representation.

    There are different flavours of DGGS. H3 is being the most popular but there are notable others such as Healpix or S2.

    • Hexagons are good for an accurate local representation and just generally look really nice. https://h3geo.org/
    • Healpix approach has a few advantages that I’m not sufficiently involved with the subject to understand. But there is a lot of good information here should you be interested.

    If your interested to follow up Id recommend starting with the ESRI blog which illustates with the following graphics how you might aggregate and visualise data using a DGGS.

    DuckDB-Spatial: Supercharged Geospatial SQL!

    As always at conferences there are many new things. If you can just process one or two ideas then it can be worth it. DuckDB for me is one of those ideas. In HxDR we use Postgres and PostGIS, DuckDB is a PostGIS inspired DB. It processes faster, fails gracefully, and appears to include a much stronger connection to the domain. Its supported by both AWS and Azure and is worthy of spike in the future. https://duckdb.org/

    Lonboard, GeoArrow, GeoParquet

    https://developmentseed.org/lonboard/latest/
    https://developmentseed.org/lonboard/latest/assets/hero-animated.gif

    Geopython introduced me to a project called Lonboard. Its a visualisation project which has connections to GeoParquet and GeoArrow.

    GeoArrow is a new data protocol which is speeding up the data delivery. Its doing this by pointing at data storage locations using a data sharing protocol. GeoArrow is the concept that the binaries of the data should be accessible directly. https://geoarrow.org/format.html

    I had a super discount airBnB which turned out to be the best apartment I’ve every stayed in.

    Thanks for following as I continue with my Geospatial musings,

    Lucas

  • Rivers of Power  by Laurence C. Smith.

    Rivers of Power by Laurence C. Smith.

    I finished a great book and would like to share; Rivers of Power – How a Natural Force Raised Kingdoms, Destroyed Civilizations, and Shapes Our World. It explains the world from the physical geography of Rivers, covering among other things:

    • The first sustained attempt at data science by the Egyptians
    • The Roman system of water measurement
    • Technical innovations in Canal design and the future of City development

    For each River I’m using a stream-lit app on top of the pretty maps project. https://prettymapp.streamlit.app/ its a python project to extract open street maps with a fancy FE.

    Egyptian Pharaohs Censor the Nile Records

    Data scientists usually appreciate the importance of good data collection. Prior to Rivers of Power I had no knowledge of the Egyptian Nile Flood records. The first systematic environmental monitoring project we’re currently aware of. An archaeological artifact called the Palermo Stone was discovered which used to record the height of the Niles flooding. The height of flood water was then used to predict the futures harvest. If the waters reached level 12 then it predicted starvation whereas if the waters rose all the way to 15 there would be widespread prosperity. The measurement devices show signs of having been hidden. This is likely because predicitive metrics about coming starvation was sensitive data. Making this the first recorded instance of environmental data censorship too.

    The Nile flows through the heart of Cairo and is 6650km long and has a discharge rate of 2,830 m³/s. Its either the longest, or second longest river in the world; the Amazon has recently been declared slightly longer…

    Ransoming Water

    Rivers of Power was my first introduction to the concept of water towers in geographic terms. Countries with large bodies of water such as Lesotu, Ethopia, Tibet, Switzerland are referred to as water towers. They have no access to the ocean for trade but they have the ability to capture water for dam projects. Downstream neighbors are often wealthier due to maritime trade opportunities. This causes some geopolitical situations such as the one unfolding at the moment between Egypt and Ethopia. Poorer countries damming water sources of wealthier states. Follow up on the Grand Ethiopian Renaissance Dam project for a great modern example. The book details how rich counties are currently de-constructing dams for environmental reasons. While developing countries are still constructing new ones.

    Switzerland is a “Water Tower” feeding the Rhine. The Rhine Falls is particularly spectacular if your ever on the Swiss German boarder.

    The River forms a natural barrier between France, Germany and Switzerland. It flows through Basel, and forms the boarder in the city between the three countries. It give the place a dynamism. The Rhine Length is 1233km and it has a discharge rate of 2,000 m³/s.

    Rivers, Health and the Thames

    Humanity exists between the ocean and the land as many of the great cities of the world boarder the sea. But it does seem to be much more accurate to say that humans are a river species. Only 5% of the world population lives away from Rivers and only 20% of the population next to the ocean. Much of humanity is moving to cities and river fronts within these cities are getting more development. And with this move more protection from extreme flooding events. The Thames River is an almost perfect example of this. Some of the most expensive real estate in England is on or next to this river. There are a serious number of developments are in place to control and develop old industrial sites along it. The Thames is only 346 km long, with a discharge rate of 65.8 m³/s.

    The book dives into factors such as Green Exercise being more effective for wellbeing. I like one of the phrases particularly “nature deficit disorder”, which basically says go out side and spend time near running water to help mental wellbeing. There is something about rivers that help restore attention more effectively than other types of rest and recovery. The soft focus required to be in natural aquatic environments as opposed to cities with more active distractions.

    Bridgewater Canal, Manchester, 1767

    During the industrial revolution the Bridgewater Canal was constructed by the engineer James Brindley. The project was designed to transport coal from Worsley mines to Manchester. This slashed coal prices by half fueled the city’s industrial boom.

    The Bridgewater Canal is for recreational usage and passes directly next to Old Trafford. The home of the Manchester united Football Club. This rendering of Manchester also came out very well on Pretty Maps so I really had to include it. The canal is 65km.

    This quickening of large goods shipments has continued to today. Navigable waterways are still the primary and cheapest means of goods transportation. Waterway engineering projects are active all over the world from Turkey to Mexico. Sues Canal bring in substantial capital and media attention. While navigable waterways are reliable predictors of economic prosperity. The Mississippi region is a source of great capital generation for the United States. Low transportation costs and navigable, its the most prized river system in the world.

    Roman Waterways Measurement and Le Charente

    The Romans were known for their mastery of waterways. Yet they had basic misunderstandings about aquatic engineering. They believed that widening a watercourse would increase discharge rate. This is a fundamental misunderstanding of hydrological measurement. River discharge is the volume of water passing a specific point over time which is measured in cubic meters per second (m³/s). This depends on both the cross-sectional area and the velocity of the water, not the width.

    Thanks for reading. Don’t be a Roman and measure your rivers by the 10 most powerful rivers by discharge:

    1. Amazon
    2. Ganges
    3. Congo
    4. Orinoco
    5. Yangtze
    6. Negro
    7. Madeira
    8. Rio de la Plata
    9. Brahmaputra
    10. Mississippi
  • Spatial SQL

    Spatial SQL

    SQL is a standard language for interacting with DB’s and I don’t think we discuss it enough in the Spatial community. It’s a de facto language for interacting with DB’s which is platform agnostic. There are some great courses which I’ve dipped in and out of but the best is to just use the tools to do something useful. 

    One of the things I’ve advoced for in my current position is that data be stored whenever possible in a database with a spatial capability, my favourite being Postgres + PostGIS. I’ve also seen a great spatial developer working with Microsoft SQL server successfully. However recently a lot of people trying to make newer flavours of DB such as Mongo and Dynamo work spatially with less success.

    Here you can see I’ve been asked to populate a Postgres instance with geographic extensions to valid vertical references systems in the United States and Europe. PostGIS is incredible in this circumstance because I can easily connect it to QGIS and work either in SQL or through a UI, or a bit of both.

    Creating a new table

    Creating a well-structured table is a fundamental skill that every SQL practitioner should master. A well designed table not only ensures efficient data storage but also plays a crucial role in optimizing database performance. In this guide, we will delve into the process of creating a table in SQL with a professional touch, incorporating essential elements such as CREATE SEQUENCE and COMMENT ON TABLE.

    SQL tables serve as the foundational building blocks for organizing and storing data in a relational database. The CREATE TABLE statement is the key command that enables developers to define the structure of a table, specifying columns, data types, and constraints. However, for a truly comprehensive approach, we will explore two additional features that elevate the sophistication of your SQL table: CREATE SEQUENCE and COMMENT ON TABLE.

    CREATE SEQUENCE is a powerful SQL feature that allows for the generation of unique numerical identifiers for a column automatically. By employing this, you can enhance data integrity and simplify the process of creating surrogate keys or unique identifiers within your table.

    Furthermore, COMMENT ON TABLE is a valuable yet often overlooked functionality. It enables developers to add descriptive comments or annotations to the entire table, providing insights into its purpose, usage, or any other relevant information. This proves especially useful for team collaboration and documentation, making it easier for future developers to understand the nuances of the table.

    Errors inserting data

    Once you have a new table I’ve hit an error while working with QGIS and PostGIS. Although you have a spatial DB, it seems to not accept geometry until at least one feature already exists. So here’s where SQL can help in data loading.

    INSERT INTO public.my_secret_table(id,name,geometry)

    In order to get the geometry I used the QGIS field calculator to build it using the inbuilt commands, generating a new field within the attribute table and then copying out the below geometry. 

    VALUES (1, 'EPSG:5703', ST_GeogFromText(MultiPolygon(((-144.61657396 82.13599465, -47.79067056 85.53957186 -46.49965852 0.8022358, -143.91238557 1.27169473, -144.61657396 82.13599465)))))

    One thing that has tripped me up a few times now is the distinction between geometry and geography on a PostGIS DB. Here is the article I used to learn about this, but in essence Geometry is faster because calculations are not in decimal degrees with a bunch of extra maths.  

    VALUES (1, 'EPSG:5703', ST_GeogFromText('MULTIPOLYGON(((-144.61657396 82.13599465, -47.79067056 85.53957186, -46.49965852 0.8022358, -143.91238557 1.27169473, -144.61657396 82.13599465)))'))

    While making a few updates from the UI I did a few typos which for some reason I couldn’t correct through the UI, so I simply went back to SQL commands for the updates, like this: 

    SET name = 'EPSG:5703'
    WHERE id = 2;
    UPDATE public.my_secret_table

    And deleted the data like this:

    DELETE FROM public.my_secret_table
    WHERE id = 1;

    And finally if you would like to work primarily in SQL then PgAdmin can visualise your geometry column which is a nice little touch. 

    So SQL is useful, it is not a specilized language like C, it’s designed to be a more human readable language which includes a binary representation of geometry. Hope this helps, Lucas.

  • LAS srs and assignment in PDAL

    LAS srs and assignment in PDAL

    One of the most popular, or even the most popular formats of 3d point cloud data is LAS. I’ve written about it before here, but in this example let’s take it further, let’s write some data. I’ve got some data with questionable srs info to start and I’m trying to work out exactly the issue.

    To export the .las we run a pipeline to convert a web optimized format for point cloud data and compress it down into .las. Once we open this file in regular tooling the file is valid, but the location is off. 

    The file is not corrupted

    I tried forcing an SRS to confirm the file was not corrupted and it worked. I found a meta data file which included a reference to the srs, which just seemed to suggest that there was no data. 

    “srs”: {}

    So the next step is to assign an SRS.

    LASPY and PROJ

    According to my understanding the python project laspy in conjunction with pyproj should be able to write the srs to the file directly; certainly laspy can modify and write the point cloud itself. After 2 ish hours of experimentation, saving subsets of point cloud files etc with laspy I found the docs on how to save a pyproj srs just too confusing. The project has a method to handle this operation, it’s just not clear to me how. But for the purposes of future testing I could follow the tutorials to write out point cloud histograms, which could come in handy.

    PDAL

    So I’ve given up in frustration and moved onto PDAL. As its name might suggest to you, take its inspiration from GDAL, the Geospatial Data Abstraction Library. PDAL being the Point Cloud Data Abstraction Library. PDAL has a nice documentation page but is a little intimidating as a C/C++ library which breaks the process of working with data into pipelines.

    Once I’d invested the time necessary to get PDAL fully installed I was able to use it as I might GDAL; running basic inspections on data and outputting that to the terminal. In this way I could even more clearly identify that the srs was missing from the target las.

    pdal info --metadata house_32126.las --> "spatialreference": ""

    Here I could even more clearly state that the SRS had not been set, not just by reading the metadata but also by inspecting the file itself.

    PDAL Pipelines

    After configuring a basic pipeline to assign the srs to the las file I managed to get the appropriate details filled out perfectly, with some small file size reduction in addition. Unfortunately the data is still located in the incorrect place so my current assumption is that the coordinates are either incorrect or belong to either EPSG:32610 or EPSG:32126.

    I’ve attached a simple PDAL json pipeline to assign the intended epsg code.

    [{"type": "readers.las",
    "filename": "house_32126.las"},
    {"type": "writers.las",
    "a_srs": "EPSG:32610",
    "filename": "house_32610_pdal.las"
    }]

    Just get PDAL installed and run, with all the files on the same level.

    pdal pipeline -i pdal_pipeline.json

    The result of this PDAL pipeline successfully assigned the appropriate coordinate system to the data which is then visible and recognised in QGIS and is located correctly in the state of Oregon. So that’s how you analysis the crs and assign one using PDAL.

  • HackZurich 23

    HackZurich 23

    This year’s HackZurich challenge was to take information about a train shunting station in which locomotives and wagons come together on a network and create some kind of routing solution.

    Locilizing the information

    While we have .las, models, drone photography and .obj files, in order to solve a routing problem you have to build a network. Diggiging through the fancy data formats we found some .pdf’s and .dwg files which appear as if they may be geolocated but ultimately we wanted to extract a network.

    I managed to pull out quite a bit of info from a CAD file, but making that usable was another challenge.

    Modifying CAD plans to a Geospatial format

    Manually digitizing a small network for a proof of concept actually turned out to be easier than forcing Fusion360 to generate a Geospatial compatible .dxf file with the attributes in the right place.

    After some tests we ended up with a pretty stange network with node edges, nodes and edges. 

    We did this to match the network which was essentially available to the competing teams on the Siemens challenge though a real time api call; the only critical missing information was the geospatial location of the various references highlighted in the code block below.

    {
          "id": "GLIS.WEICHE.236",
          "neighbors": [
            { "neighborID": "GLIS.WEICHE.235", "side": "fixed", "neighborsSide": "left" },
            { "neighborID": "GLIS.GLEIS.GS6.1", "side": "right", "neighborsSide": "a" },
            { "neighborID": "GLIS.WEICHE.237", "side": "left", "neighborsSide": "fixed" }
          ]
        },
        {
          "id": "GLIS.WEICHE.237",
          "neighbors": [
            { "neighborID": "GLIS.WEICHE.236", "side": "fixed", "neighborsSide": "left" },
            { "neighborID": "GLIS.GLEIS.GS7.1", "side": "right", "neighborsSide": "a" },
            { "neighborID": "GLIS.GLEIS.902_", "side": "left", "neighborsSide": "a" }
          ]
        },
        {
          "id": "GLIS.WEICHE.238",
          "neighbors": [
            { "neighborID": "GLIS.GLEIS.902_", "side": "fixed", "neighborsSide": "b" },
            { "neighborID": "GLIS.WEICHE.239", "side": "right", "neighborsSide": "fixed" },
            { "neighborID": "GLIS.WEICHE.240", "side": "left", "neighborsSide": "fixed" }
          ]
        },

    Of course a manual process for such a thing is not ideal but in a time bound exercise like this it was kind of ok to literally lie on the schematic and fix it up for a working concept at a smaller scale.

    Visuals

    Taking this information into a 2D view with routing information was done by converting from the Swiss national grid to 4326 and then displaying with cesium, which can handle geospatial data in geojson format.

    Its a hackathon so another member of our team put AI into the prototype because why not.

    And volia a working prototype for presentations, video editing and un-necessarily epic into and outro music…

    Thanks team HackZurich and everyone at the Siemens booth,

    Lucas

    HackZurich Submission links:

  • Drone Licenses

    Drone Licenses

    Regulations almost always lag behind technological innovations, and that is certainly true in the Drone space. Finally enough airports have been closed in Europe to make it necessary for the EU, Iceland, Norway, Switzerland and the UK to make some drone licensing rules.

    While the UK license is not valid in Europe the Swiss and EU licenses can be used interchangeably; this is managed by the European Union Aviation Safety Agency of which Switzerland is a member.

    The Swiss training, licensing and exam info can be found here at www.uas.gate.bazl.admin.ch however the quality is not great and at the moment you have to take the exam in French, German or Italian; also you have to be able to understand the training material in one of those languages.

    The alternative to this is to take the European Exam and training which is available in English, in all levels A1, A2 and A3 in the Open Category; however, you must pay. But if you don’t mind paying for a license from between 99 – 199 EUROS then this is probably the faster and more comprehensive approach. 

    Study Material For the A1 & A3 Categories

    The exams are broken up, A1 & A3 are taken together as the basic exam and then A2 in addition which focuses on more advanced topics and the ability to fly in an urban environment. 

    Categories of drone flight are divided into 3 categories, A1, A2, A3.

    • <25kg is the open category
    • Specific Category such as below line of sight operations
    • Specialized Category is for the most risky operations such as transporting people or substances

    And then in addition you need to consider the classification of the drone by weight.

    • C0 < 250g drone
    • C1 < 900g drone
    • C2 < 4kg drone
    • C3 < 25kg drone
    • C4 < 25kg aircraft

    Flight zoning 

    Somewhat unfortunately it’s basically impossible in Switzerland now to fly anywhere near where you live. You can find the map of flight zones here. 

    And you can find the aeronautical charts here which show off a lot of the obstacles such as wind turbines.

    Control Frequencies

    There are also some brief study materials on radio wave frequencies as a precursor to understanding when the control of the UAC will be affected by environmental conditions. 

    Insurance

    Article 4 of the regulations state “Aircraft operators shall be insured in accordance with this Regulation as regards their aviation specific liability in respect of their parties” which basically means you need insurance up to one million euros.

    There is a big difference between a professional and an amateur here; amateurs can take out third party liability insurance including drone cover. 

    METAR

    METAR is a coded weather report used in the aviation industry that you must understand. The best explanation I found was not on the official training but rather https://metar-taf.com/explanation

    CET = UTC+1 in Winter

    CET = UTC+2 in Summer

    • 10009G19KT 060V130 means that the mean wind direction is 100°, variable between 60 and 130°. The average wind speed is 9 knots (09) with peaks up to 19 knots (KT).
    • VRB01KT – Wind direction not given
    • 00000KT – there is no information
    • /////KT – direction and speed cannot be determined

    In order to be able to do this you need to be able to read a METAR 

    Random Note

    And as a totally random side note during this exam I came across details on atmospheric composition which I didin’t know despite being concerned about climate change. I honestly had no idea that Co2 was only 0.03% of the atmosphere. Feel free to leave a comment if this is news to you also, thanks for reading.

  • Building with GraphQL

    Building with GraphQL

    In order to create and maintain any kind of system you need to be able to perform operations on your data. GraphQL solves some very specific problems to do with how you perform CRUD operations. However it is pretty important to understand some of the limitations with restful routing before graphQL makes any sense at all. 

    At its best, rest provides a clear architectural style for Client server interations, and you can read more about that here.

    RESTFUL ISSUES

    Suppose you want to make a POST, PUT, GET or DELETE request to a standard blog application such as POST /<blog_name> or UPDATE /<blog_name>/:id it’s simple and almost English. 

    Then extend the example to start with a user such as POST /<username>/:id/<blog_name>/ and we want to make the application posts start from a user and then post a new blog article.

    The issue becomes that in order to do what you need to in practice, you need to put more and more information together which means one of two strategies. You either break restful conventions by making endpoints such as /user/23/friends_with_companies_and_positions/ which is not good because you’ve squished too much information together or a lot of endpoints which is painful to maintain.

    In summary restful routing with highly relational data can become overwhelming. It also causes endless engineering discussions about the api strategy.

    GRAPHQL

    GraphQL is rather based on a graph in which you traverse the graph to collect the necessary information in response to a request. For example you might start with a user, expand that to users, then collect all companies associated with users. 

    It’s important to remember that GraphQL is just one small component part of an app. When traffic hits the app many things happen, and one of the things is a check to see if this is a graphql call, and if so then it will be forwarded onto GraphQL where it can make GraphQL type interactions.

    SCHEMA

    when using graphQL you must have a schema which informs GraphQL about your data’s structure and relationships. This is named schema.graphql and it looks something like this.

    type Book { title: String author: Author } 
    type Author { name: String books: [Book] }

    Using a schema to full effect you should design it around the operations of your clients. There are basically just 3 types of query that they can use

    • Query – for retrieving data
    • Mutation – for modifying data
    • Subscription – for receiving notifications about data

    As an application grows so does the schema, and making sure those changes dont break current implementations is the real art of using GraphQL.

    Fragments

    A GraphQL fragment is a reusable piece of a GraphQL query that allows you to define a set of fields that you want to retrieve from a GraphQL server. Fragments help you avoid redundancy in your queries and make your code more organized and maintainable. They are an essential part of GraphQL’s query composition and reuse capabilities.

    Here’s an example of how you might use a GraphQL fragment:

    
    # Define a fragment to retrieve basic user information
    fragment UserInfo on User {
      id
      username
      email
    }
    
    # Use the UserInfo fragment in a query
    query {
      currentUser {
        ...UserInfo
      }
    }
    
    # Use the same UserInfo fragment in another query
    query {
      getUser(id: 123) {
        ...UserInfo
      }
    }

    In this example, the `UserInfo` fragment is defined with common fields for a user. It is then included in both queries (`currentUser` and `getUser`) to retrieve the same set of user information without duplicating the field selection.

    And finally here is the practical application of a GraphQL query to a geospatial app I’m working on.

    Thanks for reading,

    Lucas

  • Bitcoin basics and Web3.0

    Bitcoin basics and Web3.0

    You can think of Bitcoin as both a currency and a bank which exists on the web and was invented / discovered by Satoshi Nakamoto in 2009. Since then governments have come to grips with the idea of a digital currency and it seems that with the collection of taxes and official regulation, crypto and particularly bitcoin (BTC) are here to stay at least for the moment. 

    You can find a tracker for the current market value of BTC here.

    Technical Fundamentals of a bitcoin block

    A ledger is a record of a transaction. Regular banks use them to take note of credit and debts on their books. In the block chain space that ledger is a complete ledger of all bitcoin (or other coin) transactions at a given time which is sealed with a valid hash key; this is a block. Then the next iteration begins on the next version of the ledger, which is added to the chain of legers or chain of blocks. The mechanism is kind of like a gitrepo that cannot be rebased or a linked list in which the memory cannot be changed.

    Technically speaking Crypto is just software written in C++ with all the protocols for sending, receiving and discovering Bitcoins. 

    What is an exchange

    In order to acquire bitcoin, you must either buy it from someone or discover one (mining)

    A cryptocurrency exchange is a platform that facilitates the buying, selling, and trading of various cryptocurrencies. It functions much like a traditional stock exchange, where users can trade digital currencies for other cryptocurrencies or traditional fiat currencies like the US Dollar, Euro, etc. These exchanges provide a marketplace where users can place orders to buy or sell cryptocurrencies at specific prices.

    Coinbase is one of the best known exchanges, and although their fees are a little higher they have done a great job of making the crypto markets more accessible to everyday folks. To buy crypto you just signup with coinbase or a competitor, load cash to the platform, select the coin to buy and make an order. You will then have a crypto balance rather than a USD or EURO balance. However, now you just have a balance in Coinbase, whereas you probably want to have those coins in your wallet…

    What is a Wallet

    A wallet is a piece of software and / or software that lets you store crypto more securely. Ownership of something can be a little tricky. Most countries in the world have tried at some point to confiscate property, and private ownership is illegal in a lot of countries. To give a very specific case, the Gold Reserve Act of January 30, 1934 required all US citizens to surrender their holdings. 

    That makes all ownership of value not exactly guaranteed; with the exception of Bitcoin, in principle. The ownership of bitcoin is an intellectual idea, and to own it you must have information about the wallet in which it is held. This is a 12 wording combination. Of course this information can be stolen but it is possible to hold this in your head and it would then become the only thing which is possible to truely own; at least that’s what crypto enthusiasts would like to say. Here is an example of a wallet being recovered via the coded phrases from MetaMask. You can use an app as your wallet either on mobile or desktop or you can opt for a more secure USB device; but fundamentally each wallet has a 256 character address which is all you need to send money to someone in this form.

    Types of Wallet

    Some popular wallets include the Coinbase wallet, Tezor and Exodus but there are distinctions between wallet and you might want a couple. A web3 wallet, one which is available in your browser as a chrome extension such as the coinbase web3 or MetaMask is typically known as a hotwallet, used for daily transactions, is convenient but not very secure. 

    An alternative is a software based wallet such as Exodus which is a little safer.

    Then a cold wallet which takes your crypto offline and can include the need for you to physically press a button in order to allow transactions to take place thus preventing hacking attempts.

    Transfering coins from Coinbase to a Wallet

    In order for a transfer to occur you need to utilize a network to record the translation in a ledger and to then have that translation added to the blockchain. To understand this is quite tricky, and I’m not there yet; but then again I’m sure understanding the SWIFT network is also pretty tricky. Using this network is pretty easy in practice, you have an address and coinbase can send and receive with a scan code…. That’s pretty much it.

    As your transaction must be recorded on the blockchain you have to pay a fee for this transaction, this is called a gasfees. 

    Adventures into Web3.0 and Finance

    When you dive into the world of Crypto it pretty tempting to understand other components. First of all what is Crypto mining and why is everyone on Youtube shorts talking about it?

    What is Mining

    Every 10 minutes the bitcoin network releases an ever decreasing number of bitcoins as a programmatic reward for miners solving computation problems. There is a limited projected number of BTC possible at 21 million and 19 million are currently in existence. This will end in 2140 when there will be no more BTC to be found.

    This approximate 10 minute interval is the mechanism which keeps crypto timestamps and prevents double spend of crypto coins. As BTC are lost over time this means the value will inevitably rise assuming the same adoption rate. And so some people choose to set up mining rigs to find these remaining 2 million BTC, however as the supply is predetermined the more people mining for coins just means that mining is less profitable.

    Inflation

    Inflation is the rate at which your money loses its value, effectively; and most reliable  currencies target a 2% inflation rate. Here for example is the Swiss inflation rate for the CHF. 

    As far as I understand, an inflation rate of 2% is generally accepted as a good thing to encourage people to put their cash to work in a productive manner; however, the flip side of this argument is that it encourages consumerism. In either case I’m no economist, for me the CHF is just fine but I have opted to buy assets rather than keep cash around even in CHF. And to speak candidly I do not quite understand why the prices for everyday things should be on the rise at the same time that society has embraced technology so heavily. We’re trying to make everything easier and easier to produce, so prices should go down, not up.

    Certainly hyperinflation of Argentina is not a currency anyone outside of Argentina would be willing to accept, I believe they are currently at 104% inflation, meaning their currency loses half its value officially each year; and of course the official number is on the low side. My understanding is that in Argentina you aim to spend your salary as soon as you receive it. You should also, in theory, have to re-negotiate your salary every year just to keep up.

    NFT’s 

    If you’re in the Crypto space you’ll soon hear about NFT’s. They are tokens to represent the official copies of artwork. However I’m not sure they are going to function that way for long, my guess is that tickets of all kinds, airline, conferences, music festivals will soon become NFT tokens as they allow for the creator of the token to take a % profit of any transaction done ontop of that token.

    I wanted to try this out for real so I created a piece of nostalgic art, built up some hype, minted it as an NFT and auctioned it on Opensea. Please feel free to check out the auction but essentially I sold the piece faster than I intended for the equivalent at the time of 200 Euros or 0.02ETH.  

    Assets Investment Vehicles

    The discovery is that there is the possibility to create a valid functional monetary unit through a mathematical operation and long chain cryptographic process and protocol. This can only be discovered once as its the discovery of a functional combination. Others can innovate ontop of this, but the breakthrough is just less impressive; I wouldn’t personally invest in additional coins, although many do and many are very successful with this. I’m looking for a practical asset class to supplement the more classical savings approach. 

    Large investment funds seem as if they are now adding BTC to their investment portfolios and for the first time BTC is taking over 50% of the crypto market. More forward thinking countries are legitimizing the currency by legislating for it and even naming it as the official currency. 

    Of course time will tell and I’m not recommending anything to anyone especially other than myself but if you have any thoughts then I’d love to hear them. Also please let me know if I’ve made any mistakes here.

    In summary BTC has been around for 13 years now and it’s been attacked quite aggressively, as far as I can tell it withstood the test of time. I was sceptical initially but I think it’s probably time to reconsider. 

  • Quick GDAL

    Quick GDAL

    WHAT IS GDAL / OGR

    GDAL or the Geospatial Data Abstraction Library is awesome. Its one of the foundations of geospatial software and mastering it can really help with daily work processes.

    Originally this thing was two separate libraries, GDAL for Raster Data manipulation, ORG for vector manipulation. This separation between the two libraries still exists but the two are now bundled together.

    A lot of Geospatial tools use GDAL under the hood, so you’ve probably used it without knowing. A good place to find them is in QGIS and you will see a lot of the tools are GDAL commands with a UI over the top. And that is because fundamentally GDAL/OGR is a command line tool, which is super helpful for automating tasks. But if your only doing the task once then perhaps it’s best to just complete it in QGIS with the help of a UI.

    So without further ado, here are some of the things I’ve been using GDAL for recently.

    Case 1: It can be alternative to export data when other tools are failing

    I’ve hit the use case where for some reason more advanced (by that I mean something with a UI) fails to export large data without losing data. When this happens I used to revert to Python but now I go directly to the ogr2ogr command (for vector).

    Here are the relevant docs. And it’s also useful when you want to make some kind of selection export such as:

    ogr2ogr -where "GP_RTP=1 and Shape_Leng<1" eu_motorways_tmp.gpkg GRIP4_Region4_vector_shp/GRIP4_region4.shp

    Case 2: Using SQL embedded into a GDAL command

    Fix this case, make a git push and update the script.

    ogr2ogr "$COVERAGE_OUTPUT" "$COVERAGE_INPUT" -nlt PROMOTE_TO_MULTI -dialect sqlite -sql "SELECT ST_Union(geometry) AS geometry FROM """MyData"""" -f "ESRI Shapefile"

    This is a really efficient way to dissolve the geometry as well as being automatable.

    This complex dataset becomes the front page for our application.

    Case 3: a quick check on the diff between coordinate systems

    Often I receive data in a variety of, let’s say, exoitc formats because they are coming from the flight planning software used by local coordinate systems. Often I’m not familiar with the exact system and would like to make quick checks to see how it relates to wgs84. Rather than using some fancy tool or website you can easily do this on your terminal:

    gdaltransform -s_srs EPSG:32632 -t_srs EPSG:4326

    Case 4: automating the simplification process

    I often receive a data extract from a provider which is highly detailed, too much so to quickly load into a web server. One of the tricks I’ve come up with is running a simplification process which is available in GDAL. This one line has saved me hours and hours and hours of work.

    # Simplification Process

    echo "Simplification process begins"

    ogr2ogr $SIMPLIFICATION_DATA "$FLAT_DATA" -simplify $SIMPLIFICATION_FACTOR

    echo "Output simplified geometry file is called: "$SIMPLIFICATION_DATA""

    And there is a separate blog with more details on this here.

    Case 5: performing a quick vector to raster conversion

    Simple vector to raster (low res) in one line.

    gdal_rasterize -ts 100 100 -a_nodata -9999 -burn 1 ne_10m_admin_0_countries.shp admin.tif

    Warning: GDAL is powerful and while experimenting I managed to output 80GB files pretty quickly without noticing, just because GDAL was putting them out so fast I assumed they were smallish; not the case.

    The same command with a roads dataset give me a quite acceptable road network in a raster format.

    gdal_rasterize -ts 100000 100000 -a_nodata -9999 -burn 1 gis.osm_roads_major_v05.shp roads.tif

    The step after this was to use some proprietary software to build the raster pyramids but it completely failed me; so this experiment ended here but it is cool to be able to quickly rasterize any feature.

    Thanks for reading folks,

    Lucas

  • AWS Cloud Practitioner

    AWS Cloud Practitioner

    I’m currently working on a cloud product which is utilizing the AWS cloud; so it was a great time to study for and take the AWS Cloud Practitioner Exam. Here are some of the details on how I prepared for the exam; but none of the details from the actual exam, just logistics.

    AWS Exam Booking

    • Here is the exam which will cost 100 USD with the course code CLF-C01.
    • I’ve taken the practice questions and have managed to get 4/10 correct. Most of my failures were around security and the marketplace.
    • There are online companies which allow you to take a more extensive practice exam but you have to pay.

    So as with any goal the first step is to be clear about the outcome and to strategize a bit and the options seem pretty simple, either:

    • Youtube videos or paid udemy courses to prepare
    • Amazon Webinars (although I didn’t initially realize this was possible)
    • Random training companies offering paid for courses
    • Or the New AWS training game

    I had a little experiment with the AWS training game because it seemed like a new way to approach learning technical concepts through a VR environment. Essentially you walk around a VR world taking 12 core training modules and playing with the environment.

    Walking around the VR world

    You find the 12 core modules covered by walking around the VR world and get a kind of scenario for each. This includes some video lectures, a demo section in which you get access to an AWS account and then an independent challenge. Each challenge has an associated test. Some of the subjects included in lectures, follow allow exercises and diy challenges are:

    • S3 buckets
    • EC2 availability zones
    • Scaling up an EC2 Instance
    • VPC
    • DB setup in the cloud (MariaDB & DynamoDB)
    • Cloud economics – here’s a public link for that
    • IAM security
    • Peering connections
    • + a few more

    As you go through the world you get a % progress marker, which is actually quite helpful in measuring your effort to progress ratio; especially when considering exam booking.

    The VR world

    I’m absolutely impressed by the way they have gamified the learning environment. Walking around this virtual world is actually surprisingly engaging. When you don’t want to study difficult content you can just cruise around the city and do quizzes relevant to the exam or just arbitrarily modify the place with the points you get for doing the study. Perhaps that’s not the most effective study habit but it certainly does keep you engaged in the subject.

    Security Island

    In a moment of exploration I realized how well a Virtual Environment works for learning. I was losing focus and went exploring and found an entirely new area called security islandwhich has a medieval theme going on. It’s a physical (virtual space) where all the interactions are associated with security concepts which helps with memorisation.

    Also once I discovered this island I realized that you can accidentally find the other AWS trainings in this VR including:

    • Machine Learning
    • Solutions Architect
    • Serverless Developer
    • Cloud Practitioner

    Virtual Private Cloud VPC

    Security groups was the subject I struggled with most during the practice questions as the networking involved is a bit tricky when you start to think of IP addresses and access control to storage; or at least I thought so.

    Peering connections

    Configuring peering connections was a totally new experience for me. While it’s reasonably simple conceptually, actually implementing it was a challenge and a few of the acronyms flying around caused me some issues.

    • CIDR acronym CIDR (Classless Inter-Domain Routing)
    • POSIX file control systems = Portable Operating System Interface
    • Subnet…
    • NFS = Network File System

    This little diagram helped me understand subnets in a way that I just didn’t before.

    Finally taking the Exam

    I’m sorry but I cannot say much about this because it would be potentially a problem with the folks over at AWS. But I can say that I was able to take the exam over my lunch break at home but I had to reshuffle my desk so that it was possible to do that. Ultimately I passed the exam having studied a solid 12 hours using exclusively the VR game; although in retrospect I might also take an AWS webinar to explain some of the higher level, less practical questions.

  • Projection Cards

    Projection Cards

    This was a fun little voluntary project organized by Danield Huffman from somethingaboutmaps.com to produce a series of project cards.I thought it would be a nice opportunity to share some of my working notes to see if my process makes sense and just to put it out there for others to use.

    Inspiration

    First of all, here is some inspiration. This is an example from another participant that I like a lot.

    Print Formatting

    The template was provided, which was nice. 2.3 inc = 5.8cm x 3.3 = 8.3cm. I always try to start with the output format in mind because it’s too easy to get carried away with crazy amounts of data and totally forget it will be invisible.

    Originally I tried to do this in InkScape but quickly gave up and installed illustrator, sorry open source community 😦 

    QGIS data export at the right scale

    I’m using a QGIS program for the export because I want to get the graticles correct and done with a full GIS program. To do this I’ve downloaded all of the data I want to show from Natural Earth, including the graticule lines. 

    For the export I’ve used the layout function of QGIS and set the export to use the same ratio as the output card.

    Contents of the map in QGIS

    It’s a very small map and honestly, the map isn’t even the purpose of this exercise, we’re trying to show off the projection. That said, it’s pretty hard to understand a projection without reference information. With that in mind i’ve gone for a minimal amount of information

    • Land masses
    • Rivers and major lakes
    • Oceans
    • Hillshade

    And then to highlight the projection the graticule lines to really visualize the distortion, however these lines are usually pretty thin. In this case we need to thicken them up for the printing. Daniel recommended a minimum value of 0.5pt (that’s 0.17mm or 0.0069 inches).

    Working in Photoshop

    The final piece that I’ll have to do in photoshop is to use a linear burn on a hillshade to accentuate the values on the land masses. So to begin moving this to another format I’m going to be exporting these as .tif files with an associated .wld file. With the linear burn applied we get this:

    It’s a pretty cool effect. The only problem here was that I didn’t really want to have the graticule lines in photoshop but I couldn’t separate them into illustrator easily. I’m sure given more time I could work it out but it was getting a little frustrating so I just loaded it into photoshop in the end.

    Illustrator

    After creating the map in a combination of QGIS and Photoshop I exported the data as a tif and loaded that into the correct space in illustrator…. Only to realise that I’d measured badly in the original export from QGIS. See below image.

    So after re-measuring, re exporting and re photoshopping I imported it once more to illustrator and it fit the ‘Safe Area’.

    Labeling

    I think the most difficult part of any map is the labeling, and so I tend towards less than more, just because it’s a difficult subject! 

    I decided to hand label 2 things only.

    • Graticule
    • Oceans

    I did this is illustrator rather than QGIS 

    Here is the Miller card in which Daniel helped me a little by increasing font size, and fixing the halos. So it was lucky for me that I did the labels in illustrator rather than QGIS. If I had done that in QGIS then final edits would have been impossible.

    Here I thought I’d experiment with no ocean data just to change things up a little. 

    Thanks for reading and thanks to Daniel for organising you can follow his updates at https://somethingaboutmaps.wordpress.com/2022/04/08/projection-cards/

  • Why do Coordinate Systems matter

    Why do Coordinate Systems matter

    Coordinates without a CRS are meaningless numbers. So if we consider a CRS or SRS as a tech stack it has been in development for some time; and its starts with :globe_with_meridians:

    Size of the world

    I’m going to cover Coordinate Systems as thoroughly as possible and this is going to involve some history.

    As a species we’ve known the earth is round for quite some time, nearly 2500 years. A philosopher named Eratosthenes successfully measured the globe in 300bc arriving at a global circumference of 40,000 km only 75km shy of the actual number 40,075 km

    Early navigation

    During the age of discovery navigators started by hugging the coastline. Some of the early pioneers such as Vasco de Gama used this method exclusively to map the coast of Africa and ultimately cross into the Indian Ocean (1497–1499)

    These navigators knew the earth was round and had charts depicting latitude and longitude. Using the total circumference we were able to create accurate charts.

    • The problem of determining your latitude was solved using the angle of the sun to determine your position to the equator.  
    • The problem of longitude on the other hand was solved by “dead reckoning”, a good name because it killed a lot of people. This process used an approximation of wind and tidal speeds to determine distance travelled.

    As an interesting side note this also led to the foundations being laid to the modern financial system to account for the risk of sea voyages.

    Longitudinal awards

    In 1714 there was a push from the British government to improve navigation at sea. Ultimately this led to the insight by a watch maker named John Harrison that to master longitudinal navigation we need first to master time. This is why you see reference to Degrees, Minutes, Seconds and their conversion to decimal degrees when dealing with a Geographic Coordinate System. 

    One degree is equal to 60 minutes and equal to 3600 seconds:

    1° = 60′ = 3600″

    One minute is equal to 1/60 degrees:

    1′ = (1/60)° = 0.01666667°

    One second is equal to 1/3600 degrees:

    1″ = (1/3600)° = 2.77778e-4° = 0.000277778°

    From this point forwards sea navigation becomes a lot more predictable as you could accurately determine your location.

    Geographic Coordinate Systems

    The geographical coordinate system we use mostly is an ellipsoid called WGS84. Which has it’s 0° prime meridian longitude in Greenwich London. There was a time where there were many different models being actively used with different 0° locations, notably Paris, New York and Amsterdam. However due to the sheer amount of trade happening in London at the time, the standard moved towards London as the zero location. This standard was once again reinforced with the Glonass and NavStar systems using the WGS84 Coordinate System. (these are the two pioneering GPS systems from the USA and Russia)

    EPSG

    Above I used the term epsg which is a reference to the licencing authority for coordinate systems. This stands for the European Petroleum Survey Group EPSG who maintain the known and recognised coordinate systems. EPSG just refers to the group who hold the recognised database for the number of the system. Which is why you might occasionally see ESRI referring to themselves as the authority in the EPSG part of a reference system.

    The Geoid

    The WGS84 system uses an Ellipsoid rather than a Geoid. This is a generalised model of the earth, rather than a model of how the earth actually is. This is an important thing to understand when working with Z elevation values in WGS84 because it’s only a representation of the world, not an actual model.

    In reality the earth’s geoid looks a lot more like a deflated football that’s been kicked in India. 

    Just to reiterate, we work in WGS84, everything we do uses the Geographic Coordinate System which means by default that we use the Z elevation values based on the Earth Gravitational Model 2008; the current default for WGS84. This is an improvement on the 1996 EGM96, which in turn was an improvement on EGM84. 

    To introduce some more complexity the generalized geoid does not fit very well in some countries. Many countries mantin their own ellipsoid/geoids because they fit their country better. Often data needs to be in the country system for government contracts. So if you have a request to change the coordinate system to for example https://epsg.io/2056 that’s probably because the Swiss government mandates it. In some more extreme cases the government both mandates the reference system and the system itself is kept proprietary, such as is the case in Germany.

    Projected Coordinate Systems

    To continue our story of Cartography a little, one of the revolutions after the Marine Chronometer was the invention of the Mercator Projection in 1569. This allowed sailors to plot a bearing on a map and sail along a single bearing to their destination. Prior to this it was necessary for a ship to constantly recalculate their bearing as they travelled along a great circle.

    Here’s how to see the real size of countries

    The Mercator projection does not preserve area, but does preserve angle, which is why the northern countries look larger than they are but the roads look like they should.

    To get back to theory, converting a circle to a flat surface is an impossible mathematical challenge. That’s why there are so many different flavours of Projected Coordinate systems.

    This is essentially the process of taking something spherical and “projecting” it to a flat 2D surface.

    While the maths is complicated, it essentially breaks down to the preservation of angle or distance. Different models are better in different areas in the same way that different Geographic systems work better in different areas. I personally think of it more simply as this:

    Transformations

    One of the reasons for moving from a degree based system to a cartesian space is the ability to perform operations on real numbers. This is a challenge we will need to face when moving from our model space into real Geographical space as the measurement feature will no longer work. Luckily there are API’s for us to call upon from the Luciad team, but it’s still nice to know why we need those api’s rather than just doing the math ourselves. Essentially maths on a circle is hard while maths on a flat surface is, if not simple, certainly easier.

    So just to reiterate a geographic system is measured in degrees whilst a Projected System is measured in real numbers.

    Vertical Coordinate System

    In addition to Geographic and Projected coordinates, knowing your Z value is very important. This has become a lot more important since the advent of affordable drones. You have a choice when recording your Z elevation. You can use the distance relative to the geoid, which is computationally heavier. An ellipsoid, less accurate but simpler or to the take off location

    VERTCS[“WGS_1984”,DATUM[“D_WGS_1984”,SPHEROID[“WGS_1984”,6378137.0,298.257223563]],PARAMETER[“Vertical_Shift”,0.0],PARAMETER[“Direction”,1.0],UNIT[“Meter”,1.0]]

    You often see errors in photogrammetry projects because the EXIF information on the camera is incorrect.

    This leads to the possibility for compound coordinate systems which is putting a vertical, geographic and projected coordinate system details into one WKT.

    This is the information necessary to truly understand the details of the projection. 

    For further reading checkout I hate coordinates systems blog and remember, degrees for geographic systems, xy for projected systems and when it comes to the Z value, ask relative to what?

    P.s this is what the planet looks like really….

    WKID & WKT

    Another necessary acronym to be aware of is the WKID or Well known ID. This is a reference to the number of the coordinate system. For example I can refer to WGS84 by saying, EPSG:3857. This would in turn lead you to the well known text WKT, which contains the exact details of this code. 

    PROJCS[“WGS 84 / Pseudo-Mercator”,

        GEOGCS[“WGS 84”,

            DATUM[“WGS_1984”,

                SPHEROID[“WGS 84”,6378137,298.257223563,

                    AUTHORITY[“EPSG”,”7030″]],

                AUTHORITY[“EPSG”,”6326″]],

            PRIMEM[“Greenwich”,0,

                AUTHORITY[“EPSG”,”8901″]],

            UNIT[“degree”,0.0174532925199433,

                AUTHORITY[“EPSG”,”9122″]],

            AUTHORITY[“EPSG”,”4326″]],

        PROJECTION[“Mercator_1SP”],

        PARAMETER[“central_meridian”,0],

        PARAMETER[“scale_factor”,1],

        PARAMETER[“false_easting”,0],

        PARAMETER[“false_northing”,0],

        UNIT[“metre”,1,

            AUTHORITY[“EPSG”,”9001″]],

        AXIS[“X”,EAST],

        AXIS[“Y”,NORTH],

        EXTENSION[“PROJ4″,”+proj=merc +a=6378137 +b=6378137 +lat_ts=0.0 +lon_0=0.0 +x_0=0.0 +y_0=0 +k=1.0 +units=m +nadgrids=@null +wktext  +no_defs”],

        AUTHORITY[“EPSG”,”3857″]]

  • Working with Shapefiles

    Working with Shapefiles

    Shapefiles are an old file format, originally developed by ESRI, which have become a common way of working with Geospatial data; much to the chagrin of ESRI who have ever since been trying to migrate to a Geodatabase format. A shapefile is driven by is .shp extension but can contain upto 17 different files adding valuable information such as Z values. The four critical file extensions for a shapefile to function correctly are 

    • .shp 
    • .shx 
    • .prj – this contains the projection information of the shapefile
    • .dbf – this contains the data table. 

    Geometry types

    Here is the ESRI documentation on a shapefile. In essence it contains one Geometry type only, those are:

    • Points – literally 1 xy
    • Lines – two or more xy coordinates
    • Polygons – a start xy, any number of intermediate xy and a closing xy which completes the geometry. 

    There are some more advanced types such as multipart polygons and donut polygons which you could read more about here as it is a genuinely interesting subject within Geospaital data.

    5 Steps to creating a custom Shape file

    Step 1 The easiest way to create a shapefile is to download the application QGIS, working on mac, linux and windows here

    Step 2 Open up QGIS and you should see the shapefile creation dialogue

    Step 3 Create a new folder to contain all of your shapefile and save the file name. Mine here is test003

    Step 4 Click edit, add vertices and save the edits

    Step 5 Go to the file system and you will see the new shape file.

    I should say that there are many reasons why a shapefile is not the ideal data format but it is very useful for quick data edits or shaping a polygon. 

    Writing a .shp file with python

    To work with a shapefile programmatically, and outside of the ESRI ecosystem, you need to lean on a few libraries.

    • Shapely which deals with geometry operations
    • Fiona which handles the reading a writing and most terrifyingly
    • pyproj4 for all your projections and transformations

    However we can also just go ahead and use Geopandas which combines all of the above libraries into the Pandas ecosystem for data munging.

  • The simplification of our data coverage

    The simplification of our data coverage

    We have a highly detailed geojson file containing multipart polygons which represent the coverage of our data from the content program.This data needs to be fast to load but it is 320MB, all text. This is largely caused by each vertex in the polygon having a lot of xyz data. In cartographic applications we often run feature simplification to reduce the size of such data. Below in the blue line is what we have and the red line is what we want.

    Here is what our original data looks like. I’ve added some transparency to the data so that you can see the density of data in certain areas.

    This data is heavy because of the sheer number of vertices in the lines. There are a few cartographic tools to help simplify this but we ultimately decided to use the Douglas-Peucker algorithm, named after the two Canadian Cartographers who came up with it.

    Note: almost all GIS related stuff has some connection to Canada because Geospatial Computational processes first emerged there to manage the massive land areas. 

    The algorithm works recursively, halving the line while taking user input for the ideal level of detail. We have chosen to remove data within 900m. 

    (I have stolen this awesome giff from wikipidea)

    Here you can see our results. The red Polygon has been significantly reduced in size, from about 320 MB → 1.2 MB. While this might seem quite dramatic in the world of computer science I would say this is fairly normal in the Cartographic world. The reason begin that we are willingly and purposely removing information, whereas compression algorithms are most frequently looking to preserve information while storing it more effectively.

    If we examine just one of the coverage datasets to illustrate this a little more. If we take one feature prior to the simplification we are at 3276 pt vertices in the polygons vs 163. We’ve lost a lot of information, but we’ve preserved the shape at this scale.

    The final and I would say quite interesting thing to visualise here are the bounding boxes of the data. They looked a little strange at first but if you highlight the individual features you see each feature is actually a multipart polygon. This information [min_x, min_y, min_z, max_x, max_y, max_z] is used by the coverage to display. 

    I’ve written this all down in a blog because I suspect it will not come up again as a problem but it is so visual and interesting! And as we are the visual computing hub I thought it would be nice to see some of the visuals behind what we do!

  • How to work with point cloud data

    How to work with point cloud data

    Fundamentally point cloud data is just a lot of points needing to be represented in real or cartesian space. As a colleague recently pointed out to me, computer science is just moving information from one place to another. So the format in which we move data can become very significant and really understanding the schema can be enormously valuable. 

    testing area

    Side note: When I say Cartesian space I mean a non geographic projection. You could use a linear grid such as this or a projected Coordinate system

    3D data industry

    Currently the world of 3D data is a little confusing and a big reason for this is the sheer number of application; encompassing a lot of new industries, from Surveying using photogrammetry to Zbush for concept artists.

    In this post I will share a little about how to work with 3D data in a Geospatial application. Therefore we will discuss the main standard, .las files. However I should point out the .e57 files have a significant place when it comes to laser scanning. This is because you can write individual pieces of the laser scan more easily and so a lot of Surveying hardware uses and .e57 format.

    LAS structure

    A .las file is an open standard for working with point cloud data, designed by the American Society for Photogrammetry and Remote Sensing. Understanding a data storage standard and structure can help when working with or building a .las file. There are versions of the data storage specification and we are currently at version 1.4, released in 2011. 

    • The Header: This contains format info, number of points and extent of the point cloud. For my purposes the header usually contains all the info I need. 
    • Variable length records (VLR): This is where the coordinate reference system info is stored. Which is really a critical record for quality validation.
    • Point Cloud Records: This is where the individual points are stored within the format. Each point has the obligatory xyz but other data is also stored such as reflectance value or classification.
    • Extended variable length records: This was an addition to extend the VLR in the release of v1.3

    These are the fundamental chunks of a .las file so however you interact with one, these are the lowest level at which to consider the format.

    How to read a .las file

    There are three libraries you can lean on in Python to make this easier. Liblas, laspy and PDAL. I’m going to discuss a little bit about these. (sorry I’m not covering installing these. Hint: liblas is best into a virtual environment with pip, PDAL is best done with Conda because its a C++ library, laspy is easy with pip)

    Reading the Header

    Here is an example of me reading the header with liblas and there are a few differences between the version 1 and 2 of liblas, so be cautious with that. 

    # import liblas
    from liblas import file
    las_file = "/location/to/your/file/points.las"
    f = file.File(las_file, mode="r")
    #now we can print the header object!
    print(f.header)
    
    ['DeleteVLR', 'GetVLR', '__class__', '__del__', '__delattr__', '__dict__', '__dir__', '__doc__', '__eq__', '__format__', '__ge__', '__getattribute__', '__gt__', '__hash__', '__init__', '__init_subclass__', '__le__', '__len__', '__lt__', '__module__', '__ne__', '__new__', '__reduce__', '__reduce_ex__', '__repr__', '__setattr__', '__sizeof__', '__str__', '__subclasshook__', '__weakref__', 'add_vlr', 'compressed', 'count', 'data_format_id', 'data_offset', 'data_record_length', 'dataformat_id', 'date', 'delete_vlr', 'doc', 'encoding', 'file_signature', 'file_source_id', 'filesource_id', 'get_compressed', 'get_count', 'get_dataformatid', 'get_dataoffset', 'get_datarecordlength', 'get_date', 'get_filesignature', 'get_filesourceid', 'get_global_encoding', 'get_guid', 'get_headersize', 'get_majorversion', 'get_max', 'get_min', 'get_minorversion', 'get_offset', 'get_padding', 'get_pointrecordsbyreturncount', 'get_pointrecordscount', 'get_projectid', 'get_recordscount', 'get_scale', 'get_schema', 'get_softwareid', 'get_srs', 'get_systemid', 'get_version', 'get_vlr', 'get_vlrs', 'get_xml', 'global_encoding', 'guid', 'handle', 'header_length', 'header_size', 'major', 'major_version', 'max', 'min', 'minor', 'minor_version', 'num_vlrs', 'offset', 'owned', 'padding', 'point_records_count', 'point_return_count', 'project_id', 'records_count', 'return_count', 'scale', 'schema', 'set_compressed', 'set_count', 'set_dataformatid', 'set_dataoffset', 'set_date', 'set_filesourceid', 'set_global_encoding', 'set_guid', 'set_majorversion', 'set_max', 'set_min', 'set_minorversion', 'set_offset', 'set_padding', 'set_pointrecordsbyreturncount', 'set_pointrecordscount', 'set_scale', 'set_schema', 'set_softwareid', 'set_srs', 'set_systemid', 'set_version', 'set_vlrs', 'software_id', 'srs', 'system_id', 'version', 'version_major', 'version_minor', 'vlrs', 'xml']

    I’m utilising this data to make sure we have written the .las file correctly, so I need to get the date written, point record count and min max values. Virtually everything I need is contained here in the header.

    Visualising the results 

    If it’s ok to use an application then Cloud Compare is Mesh lab are both pretty good. Cloud Compare is specifically designed open source software for analysing point cloud data and its free so I’d recommend that one.

    If however you need to observe the data in 3D space programmatically we can, in principle, do that in python too. Here is a 3D display grid in cartesian space.

    And here are the results once I’ve removed the grid. A super nice clean point cloud in .las format.

    Takeaways

    Knowing about the specs of a .las and being able to visualize them has given me the ability to perform the following actions in an automated testing suite for the generation of new .las files. My application is an automated test suite and with the above context it is possible to perform the following automated tests:

    Checking placement of project
    # Read project images and create bounding box
    # Read project Header and extract .las bounding box
    # Calculate .las centroid
    # Confirm the centroid of the .las falls within the image bounding box
    # Utilise the benchmark and calculate the distance from benchmark centroid
    
    Check scale (benchmark required):
    # Calculate bounding box from las Header
    # Calculate ground area in m2
    # Compare with Benchmark and check difference
    # Raise error if over x% difference for manual inspection
    
    Check number of points (benchmark required):
    # Read number of points from Header (tested)
    # Raise error if over x% difference from benchmark for manual inspection
    

    Going forwards I’m working to create screenshot snaps to present into a visual highlight comparison. Thanks for reading, Lucas

  • The P4 Multispectral by DJI

    The P4 Multispectral by DJI

    Let’s look at some imagery collected from DJI’s P4M setup released last year. It’s being sold as a full agricultural solution. The Phantom 4 Multispectral (P4M) is shipping with one RGB camera and five narrow band sensors, including red edge and near infrared.

    (more…)
  • Automated Testing for Geospatial Applications

    Automated Testing for Geospatial Applications

    We really don’t want to break things that work, it’s super frustrating. We test to make sure that any changes we make do not break another piece of the code in our applications. Of course manual or exploratory tests are important but the dream is to get this stuff fully automated. (more…)

  • Representing Routes

    Representing Routes

    Calculating a collection route to optimally cover an area leaves you with two problems. Communication to the operators who need to be able to drive the route. Communication from the drivers who need to make changes.

    (more…)

  • Developing on a NodeMCU board with Micro Python

    Developing on a NodeMCU board with Micro Python

    I have a problem collecting rubbish and recycling saturation data. I’ve tried developing on existing ESRI apps to collect data and store it in the cloud. Its failed every time. These approaches failed because no one likes doing data entry on a phone or tablet, the guys who drive the truck. A clipboard is still more reliable than a smartphone but that cannot collect the street or time stamp. But who doesn’t like hitting an old arcade button right? I’m trying to develop a data collection method that’s manual but digital, easy and cheap. (more…)

  • GPS City

    GPS City

    Using GPS points to map out a city we get a new dimension to our maps. Using the increasingly ubiquitous GPS data available it occurs to me we have a new way to cartographically visualize infrastructure. With the combination of millions of GPS points we can build up an great view of the city. (more…)

  • Redirecting from Landfill to Recycling Services

    Redirecting from Landfill to Recycling Services

    Following on from a recent, challenging attempt to convert businesses from using landfill services to recycling. Here are the target companies.

    (more…)

  • Cheap to Environmentally Friendly

    Cheap to Environmentally Friendly

    Moving people from a cheap option to a more environmentally friendly one is tough in a commercial setting. Of the target list we managed to convert 36% of the companies with some interest in Recycling. (more…)

  • Animated vs Static Symbology

    Animated vs Static Symbology

    Compare the two images. They convey the same information, point data in a constrained Geographic space. The effect changes the whole thing.

    (more…)

  • Bus shelter after bus shelter damaged make change necessary

    Bus shelter after bus shelter damaged make change necessary

    A cities public transport network relies on having sheltered dry places to wait for the next bus; especially in a student city. Damage to this infrastructure affects the whole service. (more…)

  • Practical Pipe Inspections

    Practical Pipe Inspections

    City water networks are inspected regularly. CCTV footage helps confirm replacement scheduling. These are algorithmically generated, and are not always that reliable. Lets make some information visible.

    (more…)

  • A Pokie Question

    A Pokie Question

    Pokie machines… a questionable social vice or good fun. Regardless of your point of view, they are certainly profitable. Palmerston North City and Ashhurst spent $4,623,896.06 on them in 2017.  That’s $57.70 per person in a city of 80,000. (more…)

  • The Value of Building Footprints

    The Value of Building Footprints

    When mapping at a local scale we’re usually trying to show some sort of targeted data. This makes context important, but not so much that it takes away from the data your trying to show. (more…)

  • DEM effects

    DEM effects

    While working on a project I came across this effect. I was trying to render a height map, for both the country (more…)

  • Cleaning up the City – Graffiti

    Cleaning up the City – Graffiti

    For many people, a city’s success is judged aesthetically. Basically how pleasant an environment is it. (more…)

  • Trash Mapping with ArcPro

    Trash Mapping with ArcPro

    Looking over the invoices for council I found $##redacted (large sum of money, larger than I was ultimately aloud to publish on the infographic or my blog ) being spent on illegal dumping every year. That is an impressive sum of money for a small city. So I’m producing a map to show the annual cost and distribution of illegal dumping in the City; while telling the story of how long it take the council to pick the trash up. The data has been collected using Survey123 and Workforce and visualized in ArcGIS pro. (more…)

  • The Power of Python Toolboxes

    The Power of Python Toolboxes

    A lot of manual work can be avoided by using scripts to automate the processes. However, you inevitably end up with a lot of scripts to automate different pieces of a process. (more…)

  • Why we should 3D print buildings

    Why we should 3D print buildings

    Maps have an amazing ability to focus a discussion. As a discussions becomes bigger less people can see the map and the conversation breaks down. This is actually a really big problem.  (more…)

  • Understanding As-Builts

    Understanding As-Builts

    nullSince Roman times city planning has included the provisioning of Water. Moving forwards a couple of thousand years has taken us away from the obvious aqueduct to an immensely complex underground network of pipe’s servicing the 3 water system; Wastewater, Storm water and drinking Water.

    (more…)

  • Hexagon Reporting

    Hexagon Reporting

    We often find ourselves collecting and maintaining data. For GIS applications we use this data to answer questions about the world around us. Quite often the sheer quantity of data is overwhelming and the more we collect the harder it is to give a simple answer.

    This project is going to use Hexagon bins for the Motutapu Restoration project data. If it’s successful I will be able to report on disparate data sources by binning it into a hexagon grid across the island. This will allow me to report temporally for the islands data records using one feature service, a hexagon grid.

    (more…)

  • Responding to an Emergency with Maps

    Responding to an Emergency with Maps

    In the immediate aftermath of the Kaikoura earthquake a well practiced response went into effect. Within a few days information about the event increased exponentially as the various agencies began reporting. The tide of information was overwhelming. Organisation such as Environment Canterbury began requesting GIS support to help make the relevant information visible to people who could use it. I was lucky enough to be selected to go down and help out at the Environment Canterbury offices. Here is the process as I experienced it. (more…)

  • Buckminster Fuller and the Dymaxion projection

    Buckminster Fuller and the Dymaxion projection

    Richard Buckminster Fuller was an american inventor credited with, among other things, the Dymaxion Mapping Projection or sometimes known as the Fuller projection. It was designed to represent the earth as an island, and to emphasis humanitarian efforts across the globe, or as Buckminster Fuller would have said, “on Spaceship Earth”. I am ultimately writing about ‘Bucky’, as he became known, due to his contributions to the mapping community with the Fuller projection. The backstory is sad yet inspiring.  (more…)

  • Location Based To Do Lists

    Location Based To Do Lists

    Location based to do lists move us towards a map centric way of getting things done. Yes, they improve efficiency but the benefits goes much deeper.

    • They give our workers freedom
    • They give our managers confidence
    • They make our work more pleasant

    In short they are awesome.

    (more…)

  • Addressing the Land – A quick guide to Parcels, Titles and Addressing

    Addressing the Land – A quick guide to Parcels, Titles and Addressing

    The foundation of government is land, around which is wrapped the economy, legal system and elections. Land in New Zealand has a legal title, an address and a set area recorded by a surveyor known as a parcel. I believe there is some confusion here and it’s important to be clear when so much depends on land. Parcels and Legal Titles do not have addresses. (more…)

  • 4 Ways to Geocode in a Hurry

    4 Ways to Geocode in a Hurry

    Let’s talk about giving an address X Y coordinates. If you need to geocode something it’s probably going to be in a .csv .xlxs or .xls format. This post will talk you through adding a location and putting it on a map.

    (more…)

  • How to Ethically Harvest Online Data with Python for ESRI Integration

    How to Ethically Harvest Online Data with Python for ESRI Integration

    GeoJSON data is available from ramm.com about our cities bridges. To make it available for our engineers it needs to loaded to a local database. There are lots of metrics stored there too which can change daily, so we need to do this every night. Let’s automate it with python!

    (more…)

  • How to download Free Satellite Imagery

    How to download Free Satellite Imagery

    As a GIS professional your workflow very likely includes the acquisition of data to solve problems. In this post I’m going to be laying out a guide to Satellite geospatial data. Hopefully this will make finding data you don’t work with everyday a little easier.

    (more…)

  • A Formula for Facility Mapping

    A Formula for Facility Mapping

    This is an introductory tutorial for making facility maps using GIS. I’ll be working with ArcGIS Pro, the recipe applies equally well for QGIS or ArcMap.  (more…)

  • Four Tips for Working with Workforce

    Four Tips for Working with Workforce

    Workforce takes operations from ad hoc paper based system or a disconnected collection of GIS features to a systematic flexible and programmable to do list. At its heart there is a focus on keeping work loads simple. That means seeing only the information you need to perform the tasks relevant to you.  (more…)

  • Making Maps! ArcGIS Pro in under 5 minutes

    Making Maps! ArcGIS Pro in under 5 minutes

    3 Step Guide to Publishing Hard Copy Cartography in ArcGIS Pro

    Like almost every other GIS professional I’ve been producing hardcopy maps using the trusty workhorse of GIS, ArcMap for quite a while now; with the occasional foray on it’s slightly more hipster cousin QGIS. Over the last couple of years ESRI have been developing their cloud GIS  and a new Desktop App called ArcGIS Pro. (more…)