Backend 5 min read 854 words

Your Go code shouldn't depend on GitHub

ES
Your Go code shouldn't depend on GitHub

One of Go’s most practical ideas is that an import path is both the package name and the place to download it from. You write import "github.com/user/library" and the go command knows it has to go to GitHub and clone that repository with git. There’s no need for a central registry like npm or Packagist, and everyone knows where to open an issue.

The problem is that this convenience comes with fine print, and Iain Cambridge explains it very well in Don’t couple your Go code to GitHub: if the import path is your hosting address, your code is tied to that hosting.

The problem

Imagine your company decides to move from GitHub to GitLab. In almost any other language that means moving repositories and updating CI. In Go you also have to change the module line in every go.mod and every import in all the code that uses those libraries, because github.com/company/lib no longer points to the right place. And if someone misses something, they’ll keep downloading the old version.

When you have dozens of services and internal libraries, that change turns into a project that never finds its slot. The author tells the story of a company that ended up working with GitLab, GitHub and Azure DevOps at the same time because migrating was such a big task that they “didn’t have time”. The result: three platforms paid for in parallel because the import paths were making the decision for them.

Put like that it sounds absurd, but in the Go community it’s pretty much the norm.

The solution: your own domain

The alternative is to use a domain you own as the module name. That’s what go.uber.org or go.mongodb.org do, for example:

import "go.uber.org/zap"

That domain doesn’t host the code. It just tells the go command where it is. If the repository moves to GitLab tomorrow, you change that pointer and nobody has to touch a single line: imports and go get stay exactly the same.

flowchart LR
    A["go get go.company.com/lib"] -->|"?go-get=1"| B["go.company.com · meta go-import"]
    B -->|today| C["github.com/company/lib"]
    B -.->|tomorrow| D["gitlab.com/company/lib"]

The mechanism is simple. When the go command finds a path it doesn’t recognize, it requests the URL with ?go-get=1 appended and looks in the HTML for a <meta name="go-import"> tag that tells it which version control system to use and where the repository lives.

The configuration

The article includes the configuration the author uses for his own domain, go.iain.rocks. The nginx part tells people and the go command apart: people get redirected to the repository, and the go command gets the HTML with the tags.

server {
    server_name go.iain.rocks;
    root /var/www/go.iain.rocks;
    index index.html;

    location / {
        # If it's not the go command, redirect to the repository
        if ($args !~ go-get=1) {
            return 301 https://github.com/that-guy-iain$request_uri;
        }

        # If it's the go command (?go-get=1), serve the HTML
        try_files $uri $uri/ =404;
    }

    listen 443 ssl;
    # ... Certbot SSL configuration ...
}

And the HTML for each module is minimal:

<!DOCTYPE html>
<html>
    <head>
        <meta charset="utf-8">
        <meta name="go-import" content="go.iain.rocks/boneclone git https://github.com/that-guy-iain/boneclone">
        <meta name="go-source" content="go.iain.rocks/boneclone https://github.com/that-guy-iain/boneclone https://github.com/that-guy-iain/boneclone/tree/master{/dir} https://github.com/that-guy-iain/boneclone/blob/master{/dir}/{file}#L{line}">
    </head>
    <body>
    </body>
</html>

The go-import tag has three parts: the import prefix (go.iain.rocks/boneclone), the version control system (git) and the repository URL. go-source is optional and lets documentation tools link to specific directories, files and lines in the code.

The module’s go.mod obviously has to declare the new path:

module go.iain.rocks/boneclone

You don’t need nginx for this. Since these are static pages, any static hosting where you can point a domain will do, as long as it responds with those tags.

Things to keep in mind

The article focuses on the why and on the configuration, so here are a few things worth bearing in mind:

  • The domain becomes one more dependency. If it expires or the server goes down, direct downloads fail. For public modules, proxy.golang.org caches versions that have already been published, which softens the problem quite a bit, but you have to look after the domain just as you look after the repository.
  • For internal libraries in private repositories, you need to set GOPRIVATE (for example GOPRIVATE=go.company.com/*) so the go command doesn’t try to go through the public proxy or the checksum database, and give git access to the repositories.
  • Renaming an existing module breaks the people using it. Ideally you do this when the project starts. If you already have users, the sensible approach is to release a new version with the new path and announce it, since the old path will keep working for previous versions.

Why I think it’s good advice

It’s not a new idea, but it’s one of those that get forgotten because the default path works fine… until it doesn’t. Creating a repository on GitHub and using its URL as the module name is what almost everyone does, and for years nothing happens.

For small personal projects it’s probably not worth it. For any team with internal Go libraries, setting up a go.company.com domain takes an afternoon, and in exchange the decision of where to host the code goes back to being just that: a hosting decision, not a refactor of every repository.