<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en">
    <title>Malcolm Ramsay</title>
    <link href="https://malramsay.com/atom.xml" rel="self" type="application/atom+xml"/>
    <link href="https://malramsay.com"/>
    <generator uri="https://www.getzola.org/">Zola</generator>
    <updated>2023-08-23T00:00:00+00:00</updated>
    <id>https://malramsay.com/atom.xml</id>
    <entry xml:lang="en">
        <title>How long will it take to get through Airport Security?</title>
        <published>2023-08-23T00:00:00+00:00</published>
        <updated>2023-08-23T00:00:00+00:00</updated>
        <author>
          <name>Unknown</name>
        </author>
        <link rel="alternate" href="https://malramsay.com/talks/modelling-airport-security-queues/" type="text/html"/>
        <id>https://malramsay.com/talks/modelling-airport-security-queues/</id>
        
        <content type="html">&lt;div style=&quot;position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;&quot;&gt;
  &lt;iframe src=&quot;https:&#x2F;&#x2F;www.youtube.com&#x2F;embed&#x2F;NpdDfPwptFM?start=2637&quot;
    style=&quot;position: absolute; top: 0; left: 0; width: 100%; height: 100%; border: 0;&quot; title=&quot;youtube video&quot;
    webkitallowfullscreen mozallowfullscreen allowfullscreen&gt;
  &lt;&#x2F;iframe&gt;
&lt;&#x2F;div&gt;
&lt;p&gt;One of the biggest uncertainties when travelling is the question of
how early do I need to arrive at the airport to make it through
security in time for my flight?&lt;&#x2F;p&gt;
&lt;p&gt;This 5 min PyConAU Lightning talk walks through the process of
modelling passengers arriving at the airport
and making their way through security line.
By building upon existing behavioural research
and making a few assumptions
I find the latest time you can arrive for each flight
based on the expected length of the queue at that time.&lt;&#x2F;p&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>What&#x27;s the time Azure Synapse</title>
        <published>2022-02-07T00:00:00+00:00</published>
        <updated>2022-02-07T00:00:00+00:00</updated>
        <author>
          <name>Unknown</name>
        </author>
        <link rel="alternate" href="https://malramsay.com/post/what-is-the-time-azure-synapse/" type="text/html"/>
        <id>https://malramsay.com/post/what-is-the-time-azure-synapse/</id>
        
        <content type="html">&lt;p&gt;Azure Synapse is Microsoft&#x27;s fancy data analytics platform
designed to handle all kinds of data analytics tasks.
Unfortunately, there are significant deficiencies
in how this platform handles datetime data and timezones.
Making it incredibly difficult to use Synpase
in conjunction with other tools within the Data Science ecosystem.&lt;&#x2F;p&gt;
&lt;p&gt;Dealing with time and timezones is an incredibly difficult problem,
not at all helped by the 6 separate timezones currently in use on mainland Australia&lt;sup class=&quot;footnote-reference&quot;&gt;&lt;a href=&quot;#1&quot;&gt;1&lt;&#x2F;a&gt;&lt;&#x2F;sup&gt;.
However, despite all the complexity of dealing with time
and the corresponding timezones,
there are two main ways of storing and working with time information;&lt;&#x2F;p&gt;
&lt;ol&gt;
&lt;li&gt;Use a naive datetime that just includes the date and the time which we are going to call &lt;em&gt;Local Time&lt;&#x2F;em&gt;, or&lt;&#x2F;li&gt;
&lt;li&gt;use a timezone aware datetime that includes the date, time, and a timezone offset
which we are going to call &lt;em&gt;Instantaneous Time&lt;&#x2F;em&gt;.&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;p&gt;It is really important to note that these two types of times,
&lt;em&gt;Local Time&lt;&#x2F;em&gt; and &lt;em&gt;Instantaneous Time&lt;&#x2F;em&gt; are different representations,
and converting between them is an incredibly complex process
—talking to you daylight savings, &lt;a href=&quot;https:&#x2F;&#x2F;www.timeanddate.com&#x2F;news&#x2F;time&#x2F;samoa-dateline.html&quot;&gt;Samoa&lt;&#x2F;a&gt;, and all the other exceptions.
This conversion needs to be done carefully and deliberately,
that is, use a library specifically designed for handling the conversion
and trust that whoever spent years handling all the edge cases
knows far, far better than you.&lt;&#x2F;p&gt;
&lt;p&gt;Both of these approaches have their use cases,
&lt;em&gt;Local Time&lt;&#x2F;em&gt; is simpler and suited to collecting data locally,
for example most of our daily activities,
or monitoring cycling patterns across a bridge.
&lt;em&gt;Instantaneous Time&lt;&#x2F;em&gt; is required for data collected remotely
where we need the additional context of where the time was taken
allowing the comparison of the instant the event occurred.
As such see plenty of use in computer systems,
and also when travelling by aeroplane.
The major difference between &lt;em&gt;Local Time&lt;&#x2F;em&gt; and and &lt;em&gt;Instantaneous Time&lt;&#x2F;em&gt;
is we can convert the timezone of &lt;em&gt;Instantaneous Time&lt;&#x2F;em&gt;
without changing the instant at which the event occured,
allowing for the comparison and ordering of events across timezones.&lt;&#x2F;p&gt;
&lt;p&gt;This ability to change the timezone of &lt;em&gt;Instantaneous Time&lt;&#x2F;em&gt;
also gives us the ability to convert all the times to a standard timezone,
which makes for an easy comparison of events.
The standard timezone to use for this comparison is
Co-ordinated Universal Time (UTC), which has a timezone offset of +0.
Working with this standard timezone allows for
an important simplification of &lt;em&gt;Instantaneous Time&lt;&#x2F;em&gt;.
Rather than storing the date, time, and timezone offset,
we can store just the date, the time in the UTC timezone,
and that we have an &lt;em&gt;Instantaneous Time&lt;&#x2F;em&gt;.
This is how many programming languages handle timezones,
performing the conversion to a local time when displaying the data.
This is also the approach used by the parquet file format,
which natively supports a TIMESTAMP data type.
The parquet format is designed to hold columnar data,
that is data where all the values within a column have the same data type.
They will all be integers, strings, ...or TIMESTAMPs.
The metadata of a parquet file describes the contents,
including the name of each column and the type of data stored within it.
When the data type is TIMESTAMP,
there is a option to specify whether we are storing &lt;em&gt;Local Time&lt;&#x2F;em&gt;, or &lt;em&gt;Instantaneous Time&lt;&#x2F;em&gt;.
This flag that we set has the name &lt;code&gt;isAdjustedtoUTC&lt;&#x2F;code&gt; documenting
whether we are using the trick of converting everything to UTC described above.&lt;&#x2F;p&gt;
&lt;p&gt;The Parquet file format is the enabler of the Data Lake,
allowing software to read and write data to disk
in an efficient and portable way.
&lt;a href=&quot;https:&#x2F;&#x2F;docs.microsoft.com&#x2F;en-us&#x2F;azure&#x2F;synapse-analytics&#x2F;overview-what-is&quot;&gt;Azure Synapse&lt;&#x2F;a&gt; is Microsoft&#x27;s fancy serverless database,
specifically designed to work alongside a Data Lake.
So it makes sense that there is built in support
for working with Parquet files.&lt;&#x2F;p&gt;
&lt;p&gt;Like the Parquet file format,
Synapse has different timestamp types for &lt;em&gt;Local Time&lt;&#x2F;em&gt; and &lt;em&gt;Instantaneous Time&lt;&#x2F;em&gt;.
The &lt;code&gt;datetime2&lt;&#x2F;code&gt; type is used for &lt;em&gt;Local Time&lt;&#x2F;em&gt;,
while the &lt;code&gt;datetimeoffset&lt;&#x2F;code&gt; type is used for &lt;em&gt;Instantaneous Time&lt;&#x2F;em&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;Unfortunately the handling of datetime data within synapse is significantly lacking.
When trying to read a &lt;em&gt;Local Time&lt;&#x2F;em&gt; from a parquet file
into the &lt;code&gt;datetime2&lt;&#x2F;code&gt; type Synapse gives us this really helpful error message&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code&gt;Column &amp;#x27;datetime&amp;#x27; of type &amp;#x27;DATETIME2&amp;#x27; is not compatible with external data type
&amp;#x27;Parquet physical type: INT64&amp;#x27;, please try with &amp;#x27;BIGINT&amp;#x27;. File&amp;#x2F;External table
name: &amp;#x27;local.parquet&amp;#x27;.
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Under the covers, the TIMESTAMP within a Parquet file is an INT64 type,
counting the number of milliseconds &lt;sup class=&quot;footnote-reference&quot;&gt;&lt;a href=&quot;#1&quot;&gt;1&lt;&#x2F;a&gt;&lt;&#x2F;sup&gt; since 1 January 1970 (1970-01-01 00:00).
Even more confusing about this error message is that the documentation
for the &lt;a href=&quot;https:&#x2F;&#x2F;docs.microsoft.com&#x2F;en-us&#x2F;azure&#x2F;synapse-analytics&#x2F;sql&#x2F;develop-openrowset&quot;&gt;OPENROWSET function&lt;&#x2F;a&gt; which is how we are opening the file
clearly states that TIMESTAMP (MILLIS &#x2F; MICROS) are converted to the &lt;code&gt;datetime2&lt;&#x2F;code&gt; type.
You might think that there would be a function built into Synapse
to allow for the manual conversion of a TIMESTAMP to the internal &lt;code&gt;datetime2&lt;&#x2F;code&gt; format,
however, there is absolutely no mention of it within the documentation.
So we can&#x27;t even perform this conversion ourselves
without first implementing the function to perform that conversion.&lt;&#x2F;p&gt;
&lt;p&gt;So if reading in a &lt;em&gt;Local Time&lt;&#x2F;em&gt; doesn&#x27;t work, how about an &lt;em&gt;Instantaneous Time&lt;&#x2F;em&gt;?
When we reading the &lt;em&gt;Instantaneous Time&lt;&#x2F;em&gt; using the same approach,
rather than the cryptic error message, the file reads with no problems.
So we can just use &lt;em&gt;Instantaneous Time&lt;&#x2F;em&gt; and everything works right???&lt;&#x2F;p&gt;
&lt;p&gt;No.&lt;&#x2F;p&gt;
&lt;p&gt;Remember how I mentioned we want to be really careful about the conversion
between &lt;em&gt;Instantaneous Time&lt;&#x2F;em&gt; and &lt;em&gt;Local Time&lt;&#x2F;em&gt;,
well here Synapse is completely ignoring our &lt;em&gt;Instantaneous Time&lt;&#x2F;em&gt;
and just reading the data into its &lt;code&gt;datetime2&lt;&#x2F;code&gt; type representing &lt;em&gt;Local Time&lt;&#x2F;em&gt;.
The problem here is that other tools&lt;sup class=&quot;footnote-reference&quot;&gt;&lt;a href=&quot;#2&quot;&gt;2&lt;&#x2F;a&gt;&lt;&#x2F;sup&gt; actually pay attention to these details,
so the &#x27;quick fix&#x27; of flicking the switch to using the &lt;em&gt;Instantaneous Time&lt;&#x2F;em&gt;
is going to cause far more problems with other tools using this data.
To solve this problem properly,
we need changes at both sides of the data pipeline.
Firstly, the &lt;em&gt;Local Time&lt;&#x2F;em&gt; needs to be converted to &lt;em&gt;Instantaneous Time&lt;&#x2F;em&gt;,
assuming we know the locale in which the times were collected.
We will need this locale at the end,
so better hope it is documented somewhere
(I know that might be asking too much).
If we have no idea trying UTC and hoping for the best
at least allows us to read the file,
even if it causes problems down the track.
With the data having a locale,
one of the &lt;a href=&quot;https:&#x2F;&#x2F;arrow.apache.org&#x2F;docs&#x2F;python&#x2F;timestamps.html&quot;&gt;Apache Arrow Parquet writers&lt;&#x2F;a&gt;
will automatically convert the timezone to UTC upon writing the file.
Now when reading the file within Synapse
we need to do two conversions when selecting our timestamp column,
the first to convert Synapse&#x27;s &lt;em&gt;Local Time&lt;&#x2F;em&gt; to
an &lt;em&gt;Instantaneous Time&lt;&#x2F;em&gt; in the UTC time zone,
followed by a conversion to the local timezone we identified earlier.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;SQL&quot; class=&quot;language-SQL &quot;&gt;&lt;code class=&quot;language-SQL&quot; data-lang=&quot;SQL&quot;&gt;SELECT timestamp AT TIME ZONE &amp;#x27;UTC&amp;#x27; AT TIME ZONE &amp;#x27;Cen. Australia Standard Time&amp;#x27;
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;So how about doing the conversion manually.
Well it turns out this acutally isn&#x27;t possible
with the tools available within Synapse.
There is no inbuilt method to convert from this TIMESTAMP
to the inbuilt date types.
Additionally,
if we want to try and do this conversion manually
the &lt;code&gt;DATEADD&lt;&#x2F;code&gt; function within Synapse only supports a 32 bit signed integer,
of which the largest is 2^16-1 (~2.1 billion).
In the most optimistic case supported by the Parquet file,
that is, using milliseconds since 1970-01-01 00:00:00,
the only dates we can describe are the 25 days surrounding it.
Which really doesn&#x27;t solve the problem here.&lt;&#x2F;p&gt;
&lt;p&gt;What I find really unusual about this whole situation
is that this is definitely not a result of using a really old Parquet reader.
We can know this because there has been a change in how the
TIMESTAMP type is defined within the Parquet file.
The previous types had no notion of &lt;em&gt;Local Time&lt;&#x2F;em&gt;
and were either annotated as &lt;code&gt;TIMESTAMP_MILLIS&lt;&#x2F;code&gt; or &lt;code&gt;TIMESTAMP_MICROS&lt;&#x2F;code&gt;.
For compatibility reasons these old types are still annotated on the files
with the filetype specifications noting&lt;&#x2F;p&gt;
&lt;p&gt;Despite there is no exact corresponding ConvertedType for local timestamp semantic, in order to support forward compatibility with those libraries, which annotated their local timestamps with legacy TIMESTAMP_MICROS and TIMESTAMP_MILLIS annotation, Parquet writer implementation &lt;em&gt;must&lt;&#x2F;em&gt; annotate local timestamps with legacy annotations too, as shown below.&lt;&#x2F;p&gt;
&lt;p&gt;If the Parquet reader for Synapse was only looking at these legacy annotations
then both &lt;em&gt;Local Time&lt;&#x2F;em&gt; and &lt;em&gt;Instantaneous Time&lt;&#x2F;em&gt; would work
since they have the same type defined.
It is only in the newer file format definition
that there is anything differentiating &lt;em&gt;Local&lt;&#x2F;em&gt; and &lt;em&gt;Instantaneous&lt;&#x2F;em&gt; time.
So the new types are supported, but not properly
which in my opinion is worse than not supporting them at all.
To add even more confusion to the mix,
the type provided when trying to read the parquet timestamp
into a &lt;code&gt;datetimeoffset&lt;&#x2F;code&gt; type annotates the type
in the Parquet file as &lt;code&gt;TIMESTAMP_MILLIS&lt;&#x2F;code&gt;,
that is, using the legacy annotation.&lt;&#x2F;p&gt;
&lt;p&gt;For Synapse to be a properly fulfil a role within a Data Science toolkit
Microsoft really needs to properly support TIMESTAMPs
from the Parquet file format.
Currently, Synapse is unsuitable as part of a data pipeline
that handles datetime data.
The lack of both inbuilt support for TIMESTAMPs
and the tools for individuals to support them
severly limits the usefullness of Synapse,
and makes it difficult to recommend
as part of a broader data analytics platform.&lt;&#x2F;p&gt;
&lt;div class=&quot;footnote-definition&quot; id=&quot;1&quot;&gt;&lt;sup class=&quot;footnote-definition-label&quot;&gt;1&lt;&#x2F;sup&gt;
&lt;p&gt;During summer these are: AEST, AEDT, ACDT, ACST, AWST, and Eucla.
There are also at least 4 more in use off mainland Australia.&lt;&#x2F;p&gt;
&lt;&#x2F;div&gt;
&lt;div class=&quot;footnote-definition&quot; id=&quot;2&quot;&gt;&lt;sup class=&quot;footnote-definition-label&quot;&gt;2&lt;&#x2F;sup&gt;
&lt;p&gt;Microseconds are also a supported option. And in newer versions of the
parquet file format, Nanoseconds are also supported.&lt;&#x2F;p&gt;
&lt;&#x2F;div&gt;
&lt;div class=&quot;footnote-definition&quot; id=&quot;3&quot;&gt;&lt;sup class=&quot;footnote-definition-label&quot;&gt;3&lt;&#x2F;sup&gt;
&lt;p&gt;For example, the entire python Data Science ecosystem.&lt;&#x2F;p&gt;
&lt;&#x2F;div&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>A Workflow for Data Science</title>
        <published>2021-06-29T00:00:00+00:00</published>
        <updated>2021-06-29T00:00:00+00:00</updated>
        <author>
          <name>Unknown</name>
        </author>
        <link rel="alternate" href="https://malramsay.com/post/workflow-for-data-science/" type="text/html"/>
        <id>https://malramsay.com/post/workflow-for-data-science/</id>
        
        <content type="html">&lt;p&gt;Jupyter notebooks are a phenomenal tool
for investigating and interacting with data.
The interactive nature of the notebooks
allows for quick and effective interrogation of the data
and more generally the prototyping of an idea.
By combining interaction, the documentation, the code,
and the visualisations,
the notebooks are also a fantastic tool for communicating ideas.
However, notebooks are not the most appropriate for all the code we write,
with many there being many detractors to &lt;a href=&quot;https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=7jiPeIFXb6U&quot;&gt;using notebooks&lt;&#x2F;a&gt;
and there are definitely valid reasons for disliking the notebook.
However, most of the issues can be addressed by recognising that
the jupyter notebook is one of many tools that we have available,
and knowing when to move beyond the notebook
can help us make the most of it.&lt;&#x2F;p&gt;
&lt;p&gt;While I am using python to illustrate these steps and examples,
the same ideas apply for any language being used in the notebook
just with some slightly different specifics on the how.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;idea-development&quot;&gt;Idea Development&lt;&#x2F;h2&gt;
&lt;p&gt;The first step within the process is the exploration and development of an idea.
This is where we are just getting started looking at a problem
and we probably haven&#x27;t written any code yet.
In this phase of development we want to make full use
of the interactivity of jupyter.
This is where we can easily play around
to make sure we are using the right options for reading in our file,
trying lots of different types of analysis,
while having a record of what we have done
so that we can easily regenerate a previous state when we break something.&lt;&#x2F;p&gt;
&lt;p&gt;It is here where the interactive nature of the juptyer notebook really stands out.
It is easy to introspect the state at each step in the process,
looking at the problem in different ways to properly understand what is going on.
We also have the visualisations,
which allow us to generate nice figures representing the data,
however also includes the pretty printed outputs,
like the nicely formatted tables of dataframes,
which make working with the data easier.&lt;&#x2F;p&gt;
&lt;p&gt;The way we retain state within the notebook
is incredibly appealing,
particularly where the first step of the process is time consuming,
like loading the data from a URL,
performing a complex database query,
or doing some intermediate processing.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;organisation-of-thoughts&quot;&gt;Organisation of Thoughts&lt;&#x2F;h2&gt;
&lt;p&gt;As you start to get an understanding of the problem you are solving
and the types of analyses and steps that you will be performing.
You are going to start grouping cells together,
to make it easier to run.
Reading the dataset and the initial transformations in one cell,
another to calculate some interesting values,
and one to generate a visualisation.&lt;&#x2F;p&gt;
&lt;p&gt;This process of organising related lines of code together
is the first step towards looking at creating functions.
The functions allow you to abstract away the details
of what is going on in a particular step,
like the data loading, and focus on the important parts.&lt;&#x2F;p&gt;
&lt;p&gt;We might have the following code to read in a file&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;df = pd.read_csv(&amp;quot;data&amp;#x2F;gapminder_data.csv&amp;quot;)
df[&amp;quot;country&amp;quot;] = df[&amp;quot;country&amp;quot;].astype(&amp;quot;category&amp;quot;)
df[&amp;quot;continent&amp;quot;] = df[&amp;quot;continent&amp;quot;].astype(&amp;quot;category&amp;quot;)
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;where the datafile &lt;code&gt;gapminder_data.csv&lt;&#x2F;code&gt; is from the software carpentry training courses
and can be downloaded &lt;a href=&quot;https:&#x2F;&#x2F;malramsay.com&#x2F;post&#x2F;workflow-for-data-science&#x2F;%22https:&#x2F;&#x2F;raw.githubusercontent.com&#x2F;swcarpentry&#x2F;r-novice-gapminder&#x2F;gh-pages&#x2F;_episodes_rmd&#x2F;data&#x2F;gapminder_data.csv%22&quot;&gt;here&lt;&#x2F;a&gt;.
In loading this file we are concerned with two inputs,
firstly, the name of the file that we are loading,
and secondly the names of the fields
we want to load as categorical columns.
This results in a dataframe that is ready for us to use.
The function we might write could look something like what we have below,
where I have used type annotations to be clear about the data passed as input.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;def read_csv_dataset(filename: str, categorical_columns: List[str] = []):
  df = pd.read_csv(&amp;quot;data&amp;#x2F;gapminder_data.csv&amp;quot;)
  for column in categorical_columns:
    df[column] = df[column].astype(&amp;quot;category&amp;quot;)
  return df
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;By transforming our code snippet into a function,
we can make our intent clearer within the narrative of the notebook.
Now when we load the dataset we can use the line&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;df = read_csv_dataset(
  filename=&amp;quot;data&amp;#x2F;gapminder_data.csv&amp;quot;,
  categorical_columns=[&amp;quot;country&amp;quot;, &amp;quot;continent&amp;quot;]
)
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;which makes it clear to the reader that we want the &lt;code&gt;country&lt;&#x2F;code&gt; and &lt;code&gt;continent&lt;&#x2F;code&gt; columns to be categorical
while everything else is less important to our analysis.
This function can then be re-used within a notebook,
serving slightly different purposes each time,
loading a different file,
or allowing us to investigate whether setting columns to categorical
changes the performance of our analysis.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;solidification-of-idea&quot;&gt;Solidification of Idea&lt;&#x2F;h2&gt;
&lt;p&gt;Once we have defined a function within one notebook,
it often becomes something we want to reach for within other notebooks.
Alternatively the definition of the function
could be detracting from the story we are trying to tell within the notebook.
Once we have solidified an idea and defined a function,
we can move that function to a &lt;code&gt;.py&lt;&#x2F;code&gt; file.&lt;&#x2F;p&gt;
&lt;p&gt;By creating a file &lt;code&gt;loading.py&lt;&#x2F;code&gt; in the same directory as the jupyter notebook
with the &lt;code&gt;read_csv_dataset&lt;&#x2F;code&gt; function as its contents&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;# loading.py
import pandas as pd

def read_csv_dataset(filename: str, categorical_columns: List[str] = []):
  df = pd.read_csv(&amp;quot;data&amp;#x2F;gapminder_data.csv&amp;quot;)
  for column in categorical_columns:
    df[column] = df[column].astype(&amp;quot;category&amp;quot;)
  return df
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;we can import this function into the notebook as follows&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code&gt;from loading import read_csv_dataset
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;This works because the file loading is in the same directory as the juptyer notebook,
if it was located elsewhere, we would need to modify where jupyter looks for files to import.
I will commonly locate all my notebooks within a &lt;code&gt;notebook&lt;&#x2F;code&gt; directory of a project,
while all the python files in the &lt;code&gt;src&lt;&#x2F;code&gt; directory of a project,
a structure like below.&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code&gt;project&amp;#x2F;
├── notebook
│   └── gapminder.ipynb
└── src
    └── loading.py
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;To import the &lt;code&gt;loading.py&lt;&#x2F;code&gt; file (in this context it is called a module) from the &lt;code&gt;gapminder.ipynb&lt;&#x2F;code&gt; notebook,
the simplest method is to tell python to look in that &lt;code&gt;src&lt;&#x2F;code&gt; folder.&lt;&#x2F;p&gt;
&lt;p&gt;If we change the import to&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code&gt;import sys
sys.path.append(&amp;quot;..&amp;#x2F;src&amp;quot;)

from loading import read_csv_dataset
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;where the &lt;code&gt;sys&lt;&#x2F;code&gt; module is one of python&#x27;s built in modules
and &lt;code&gt;sys.path&lt;&#x2F;code&gt; is the list of places that python looks
for things to import.
By manually adding the relative path to the &lt;code&gt;src&lt;&#x2F;code&gt; directory
we are able to import from any files defined within it.&lt;&#x2F;p&gt;
&lt;p&gt;One disadvantage of the move to importing from a file
is that it becomes harder to update the function,
requiring a restart of the kernel to reload the module.
This can be worked around by using the &lt;code&gt;%autoreload&lt;&#x2F;code&gt; magic
built into IPython.&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code&gt;%load_ext autoreload
%autoreload 2
import sys
sys.path.append(&amp;quot;..&amp;#x2F;src&amp;quot;)

from loading import read_csv_dataset
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;which means that any time we edit the file &lt;code&gt;..&#x2F;src&#x2F;loading.py&lt;&#x2F;code&gt;,
python will reload the file and its definitions
replacing the previous version of &lt;code&gt;read_csv_dataset&lt;&#x2F;code&gt; with the new one.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;does-this-code-actually-work&quot;&gt;Does this code actually work&lt;&#x2F;h2&gt;
&lt;p&gt;Once we have moved code from a jupyter notebook to a python file
we have a range of tools available to help evaluate whether
the code we have written works.
Within the notebook this would normally be a quick process of running the code,
within a python file this can become a little harder,
particularly as the size of the file gets larger.&lt;&#x2F;p&gt;
&lt;p&gt;To help with this there are a number of tools making this process easier,&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;black&lt;&#x2F;code&gt; provides automatic code formatting so you can worry about the code
rather than how it is laid out.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;code&gt;mypy&lt;&#x2F;code&gt; provides static type checking, which can be really useful for finding
all the places &lt;code&gt;None&lt;&#x2F;code&gt; can show up in the code. Note that this only applies if
you are using type annotations.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;code&gt;flake8&lt;&#x2F;code&gt; provides a range of checks that the code works, primarily checking
for common mistakes.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;None of these tools however actually run the code,
for that we want to look towards &lt;code&gt;pytest&lt;&#x2F;code&gt;
providing tools to run the code to make sure it works.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;package-creation&quot;&gt;Package Creation&lt;&#x2F;h2&gt;
&lt;p&gt;In the vast majority of data analysis projects,
having the code in a separate file is as far as we need to go.
All the tooling is going to work in helping us keep our code working.
However, there are some cases where the functionality you are working on
is more widely applicable and needed in multiple projects.&lt;&#x2F;p&gt;
&lt;p&gt;For this we can create a package for the code,
with the package becoming a dependency of the project.
This allows for anyone to install the package
and we can use the same tools &lt;code&gt;pip&lt;&#x2F;code&gt; and&#x2F;or &lt;code&gt;conda&lt;&#x2F;code&gt;
that we use to manage the rest of our external dependencies.
For information on packaging the Python Package Authority (PyPA)
has a useful &lt;a href=&quot;https:&#x2F;&#x2F;packaging.python.org&#x2F;&quot;&gt;user guide&lt;&#x2F;a&gt;
and conda also has a &lt;a href=&quot;https:&#x2F;&#x2F;docs.conda.io&#x2F;projects&#x2F;conda-build&#x2F;en&#x2F;latest&#x2F;user-guide&#x2F;tutorials&#x2F;index.html&quot;&gt;user guide&lt;&#x2F;a&gt; for getting started.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;conclusion&quot;&gt;Conclusion&lt;&#x2F;h2&gt;
&lt;p&gt;The Jupyter notebook is the starting place for code,
it allows us to quickly iterate and try ideas out to see what works.
Additionally notebooks they are a fantastic storytelling tool,
however to make the most of them and to keep the story engaging,
we sometimes need to look beyond the notebook.
How far we go depends heavily on the type of project being undertaken
and the tools we want to use.&lt;&#x2F;p&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>Updating Query with Soft Delete for SQLAlchemy 1.4</title>
        <published>2021-03-31T00:00:00+00:00</published>
        <updated>2021-03-31T00:00:00+00:00</updated>
        <author>
          <name>Unknown</name>
        </author>
        <link rel="alternate" href="https://malramsay.com/post/updating-query-with-soft-delete/" type="text/html"/>
        <id>https://malramsay.com/post/updating-query-with-soft-delete/</id>
        
        <content type="html">&lt;p&gt;When working on a database,
there are times where we want to delete a row,
but at the same time we still want to keep around
the record that the particular row was referring to.
There can be many reasons for this,
however the basic idea is that we want the row to generally be hidden.
The Soft Delete pattern allows for the &#x27;delete&#x27; operation,
to hide the row from the external interface,
while it remains within the database.
Miguel Grinberg has a great writeup of the pattern
and implementation in SQLAlchemy on &lt;a href=&quot;https:&#x2F;&#x2F;blog.miguelgrinberg.com&#x2F;post&#x2F;implementing-the-soft-delete-pattern-with-flask-and-sqlalchemy&quot;&gt;their blog&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;This implementation is great,
however the SQLAlchemy 1.4 release changed some of the internals
meaning Miguel&#x27;s implementation no longer works,
giving the below error&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;AttributeError: &amp;#x27;QueryWithSoftDelete&amp;#x27; object has no attribute &amp;#x27;_mapper_zero&amp;#x27;
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;The fix for this issue is to replace the &lt;code&gt;_mapper_zero&lt;&#x2F;code&gt; method with the &lt;code&gt;_entity_from_pre_ent_zero&lt;&#x2F;code&gt; method, giving a &lt;code&gt;with_deleted&lt;&#x2F;code&gt; method as below.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;class QueryWithSoftDelete(BaseQuery):
    ...

    def with_deleted(self):
        return self.__class__(
            db.class_mapper(self._entity_from_pre_ent_zero().class_),
            session=db.session(),
            _with_deleted=True
        )

&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;h2 id=&quot;why-does-this-work&quot;&gt;Why does this work?&lt;&#x2F;h2&gt;
&lt;p&gt;With the fix out of the way,
we can have a look at what the &lt;code&gt;with_deleted&lt;&#x2F;code&gt; method is doing.
The entire code snippet for the &lt;code&gt;QueryWithSoftDelete&lt;&#x2F;code&gt; is reproduced below for reference.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;class QueryWithSoftDelete(BaseQuery):
    def __new__(cls, *args, **kwargs):
        obj = super(QueryWithSoftDelete, cls).__new__(cls)
        with_deleted = kwargs.pop(&amp;#x27;_with_deleted&amp;#x27;, False)
        if len(args) &amp;gt; 0:
            super(QueryWithSoftDelete, obj).__init__(*args, **kwargs)
            return obj.filter_by(deleted=False) if not with_deleted else obj
        return obj

    def __init__(self, *args, **kwargs):
        pass

    def with_deleted(self):
        return self.__class__(
            db.class_mapper(self._entity_from_pre_ent_zero().class_),
            session=db.session(),
            _with_deleted=True
        )
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;The &lt;code&gt;with_deleted&lt;&#x2F;code&gt; method provides a way to transform the Query
from one that filters out deleted values to one that does.
Now there isn&#x27;t a straightforward method for removing filters from a query,
so the approach taken here is to create a new Query,
this time setting &lt;code&gt;_with_deleted=True&lt;&#x2F;code&gt;.
The problem now becomes,
which model are we trying to query?
When we call the &lt;code&gt;with_deleted&lt;&#x2F;code&gt; method,
we are doing so from an instance of the &lt;code&gt;QueryWithSoftDelete&lt;&#x2F;code&gt; class,
so we are unable to extract the model being used through
methods like &lt;code&gt;__class__&lt;&#x2F;code&gt;, it gives us the &lt;code&gt;QueryWithSoftDelete&lt;&#x2F;code&gt; class.
So instead we can look to SQLAlchemy to get this information.
Within SQLAlchemy, every query from the database needs a &lt;code&gt;FROM&lt;&#x2F;code&gt; clause,
the table that we are going to be using to initially select our data.
This table we are selecting,
is the same one that we want to query
though this time instead of filtering the deleted items,
we will be including them.&lt;&#x2F;p&gt;
&lt;p&gt;Previously the &lt;code&gt;_mapper_zero&lt;&#x2F;code&gt; method provided that table in the &lt;code&gt;FROM&lt;&#x2F;code&gt; clause.
However the restructuring of the Query to consolidate the core and ORM components
meant the method disappeared,
completely fine given it is a &#x27;hidden&#x27; method.
We need the function that will give the equivalent behaviour,
which happens to be the much more verbose, &lt;code&gt;_entitiy_from_pre_ent_zero&lt;&#x2F;code&gt;.&lt;&#x2F;p&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>Projects</title>
        <published>2020-11-30T00:00:00+00:00</published>
        <updated>2020-11-30T00:00:00+00:00</updated>
        <author>
          <name>Unknown</name>
        </author>
        <link rel="alternate" href="https://malramsay.com/projects/" type="text/html"/>
        <id>https://malramsay.com/projects/</id>
        
        <content type="html">&lt;h2 id=&quot;web-assembly-for-visualisation&quot;&gt;&lt;a href=&quot;https:&#x2F;&#x2F;malramsay.com&#x2F;projects&#x2F;packing_wasm&quot;&gt;Web Assembly for Visualisation&lt;&#x2F;a&gt;&lt;&#x2F;h2&gt;
&lt;h2 id=&quot;pyscript-for-data-science&quot;&gt;&lt;a href=&quot;https:&#x2F;&#x2F;malramsay.com&#x2F;projects&#x2F;pyscript_demo&quot;&gt;Pyscript for Data Science&lt;&#x2F;a&gt;&lt;&#x2F;h2&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>Optimising the Learning Experience for Researchers at all Levels</title>
        <published>2020-10-22T00:00:00+00:00</published>
        <updated>2020-10-22T00:00:00+00:00</updated>
        <author>
          <name>Unknown</name>
        </author>
        <link rel="alternate" href="https://malramsay.com/talks/optimising-the-learning-experience/" type="text/html"/>
        <id>https://malramsay.com/talks/optimising-the-learning-experience/</id>
        
        <content type="html">&lt;div style=&quot;position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;&quot;&gt;
  &lt;iframe src=&quot;https:&#x2F;&#x2F;www.youtube.com&#x2F;embed&#x2F;fgVt0z-ttjI&quot;
    style=&quot;position: absolute; top: 0; left: 0; width: 100%; height: 100%; border: 0;&quot; title=&quot;youtube video&quot;
    webkitallowfullscreen mozallowfullscreen allowfullscreen&gt;
  &lt;&#x2F;iframe&gt;
&lt;&#x2F;div&gt;
&lt;p&gt;Since Intersect incorporated programming languages into our researcher training catalogue, we have seen a continuing and rapid increase in demand for courses in Python, R, MATLAB and Julia. In 2019, we trained more than 2100 researchers in programming courses alone, using the Carpentries material as a basis and complemented by Intersect’s own courses.&lt;&#x2F;p&gt;
&lt;p&gt;We notice there are three types of researchers attending our courses; those completely new to programming and hesitating about which language to use, those with programming experience but relatively new to a particular language they want to use, and those with experience in a language who want to extend their skills into libraries for data analysis and visualisation. To address this, we have established teaching pathways optimising the learning process of researchers at a level they are comfortable with and matching their skills.&lt;&#x2F;p&gt;
&lt;p&gt;At a foundational level, a new series of awareness-raising webinars introduces learners to programming concepts in preparation for interactive learning. Introductory units focus on exemplifying these concepts rather than teaching a particular language, preparing learners for any language. Finally, more advanced units build on this foundation, allowing more advanced learners to extend their skills for specific data analysis use-cases based on their research needs.&lt;&#x2F;p&gt;
&lt;p&gt;In this presentation, we explain the motivations for developing these new teaching pathways, including an analysis of course evaluations and open-ended feedback. We discuss how the Carpentries materials fit into a broader curriculum, bookended by language-agnostic awareness and introductory courses, and use-case specific advanced courses.&lt;&#x2F;p&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>Handling Exceptions in Numerical Python</title>
        <published>2020-10-19T00:00:00+00:00</published>
        <updated>2020-10-19T00:00:00+00:00</updated>
        <author>
          <name>Unknown</name>
        </author>
        <link rel="alternate" href="https://malramsay.com/post/handling-exceptions-in-python/" type="text/html"/>
        <id>https://malramsay.com/post/handling-exceptions-in-python/</id>
        
        <content type="html">&lt;p&gt;The numerical computing capabilities of the Python ecosystem are incredibly powerful,
with the &lt;a href=&quot;https:&#x2F;&#x2F;numpy.org&#x2F;&quot;&gt;Numpy&lt;&#x2F;a&gt; and &lt;a href=&quot;https:&#x2F;&#x2F;www.scipy.org&#x2F;&quot;&gt;Scipy&lt;&#x2F;a&gt;, and &lt;a href=&quot;https:&#x2F;&#x2F;pandas.pydata.org&#x2F;&quot;&gt;Pandas&lt;&#x2F;a&gt; packages
providing high performance, well tested tools.
Additionally, tools like &lt;a href=&quot;https:&#x2F;&#x2F;dask.org&#x2F;&quot;&gt;Dask&lt;&#x2F;a&gt; and &lt;a href=&quot;https:&#x2F;&#x2F;rapids.ai&#x2F;&quot;&gt;Rapids&lt;&#x2F;a&gt;
build upon these foundational packages supporting larger datasets.
All these tools are excellent 
when the data is well formed and contains the expected values.
It is when we are handling the error cases
that we can run into issues.&lt;&#x2F;p&gt;
&lt;p&gt;During my PhD,
I used scipy for the analysis of molecular dynamics simulations.
These simulations are effectively putting things in a box 
and shaking it to see what happens.
The novel component of my research
was shaking the box for a really, really, really, long time.
So long that the software I was using ran out of numbers to count. &lt;sup class=&quot;footnote-reference&quot;&gt;&lt;a href=&quot;#1&quot;&gt;1&lt;&#x2F;a&gt;&lt;&#x2F;sup&gt;
Having run the simulation,
I then needed to go though the terabytes of data generated
to calculate quantities of interest.&lt;&#x2F;p&gt;
&lt;p&gt;The huge benefit of using python and the suite of numerical tools
is the ability to quickly prototype a solution,
working through the ideas on a small scale
that can then be applied to the larger dataset.
One of the calculations I was performing
was to fit a straight line to a region of calculated values. &lt;sup class=&quot;footnote-reference&quot;&gt;&lt;a href=&quot;#2&quot;&gt;2&lt;&#x2F;a&gt;&lt;&#x2F;sup&gt;
We could come up with the following solution
for fitting line of best fit to a region of a curve,
returning both the best parameters and the standard deviation.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;import scipy

def linear(x, m, b):
    &amp;quot;&amp;quot;&amp;quot;Define the linear relationship y=mx+b.&amp;quot;&amp;quot;&amp;quot;
    return m * x + b

def fit_linear_region(x, y):
    &amp;quot;&amp;quot;&amp;quot;Fit a linear curve where the points lie within a region of 2D space.&amp;quot;&amp;quot;&amp;quot;
    # Select a region of points
    region = np.logical_and(2 &amp;lt; y, y &amp;lt; 50)
    # Perform the curve fitting only using the linear region
    popt, pcov = scipy.optimize.curve_fit(linear, x[region], y[region])
    # Calculate the standard deviation of the fit parameters
    perr = 2 * np.sqrt(np.diag(pcov))
    # Returns the optimal curve and associated error
    return popt, perr
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;When performing tests on small numbers of values,
we typically choose those behaving as we expect.
Over the course of many long simulations
nearly all possible values are expressed,
meaning this initial testing only covers a small selection 
of all the possible values.
On many occasions,
these unexpected values resulted in the code raising an exception,
unexpectedly bringing an hours long calculation to a halt.
For interactive analysis,
these exceptions are not too problematic,
you can easily fix the problem and re-run the analysis.
However when leaving the calculation running overnight,
coming back to halted execution partway through the analysis
is incredibly frustrating.
In these long running analyses,
in moving from the prototyping
to the production step of the process,
we have to consider all the places
where our code might fail
allowing us to walk away knowing a result
will be waiting for us the next morning.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;exploring-the-solution-space&quot;&gt;Exploring the Solution Space&lt;&#x2F;h2&gt;
&lt;p&gt;The problem of handling errors within code 
is faced by programmers in a wide range of disciplines.
Accordingly, there are a range of approaches
to ensure our code doesn&#x27;t unexpectedly fail.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;catch-all&quot;&gt;Catch-all&lt;&#x2F;h3&gt;
&lt;p&gt;One of the possible approaches we can take
is to capture all exceptions with a naked &lt;code&gt;except&lt;&#x2F;code&gt; block,
that is&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;try:
    popt, pcov = scipy.optimize.curve_fit(linear, x[region], y[region])
except Exception:
    return np.nan, np.nan
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;While this approach handles all possible issues that could occur,
it is also heavy handed in its application.
The general goal is to prevent unexpected errors
in a particular configuration halting the execution of our program.
However, in capturing all possible errors,
we also capture errors that can be problematic. 
For example a &lt;code&gt;KeyboardInterrupt&lt;&#x2F;code&gt;,
meaning that rather than &lt;code&gt;&amp;lt;Ctrl&amp;gt;-c&lt;&#x2F;code&gt; stopping the execution of the program,
it instead describes a curve that doesn&#x27;t fit.
This could also be problematic with a &lt;code&gt;FileNotFoundError&lt;&#x2F;code&gt;,
where instead of getting an error at the start of the program
we have to work out why our results file doesn&#x27;t exist
when trying to look through them.&lt;&#x2F;p&gt;
&lt;p&gt;While the idea of the catch-all is useful,
it is primarily delaying problems identified when the program runs
to subsequent steps in analysis,
which become more difficult to solve,
amplified by the lack of error messages
and tracebacks that python provides.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;hypothesis&quot;&gt;Hypothesis&lt;&#x2F;h3&gt;
&lt;p&gt;The main issue with the catch-all case
is that we are overly broad 
in the types of exceptions that we catch.
Rather than catching all exceptions,
what if we could determine the exceptions expected from a function call,
only catching that smaller subset.
This is similar to the approach that you would take
in an interactive type analysis
where you keep adding the exceptions that arise
to the list of those caught and handled.
An alternative to finding these exceptions manually
is through [hypothesis],
a package for enhancing test suites by
using a directed random search,
probing for values that cause problems.
The directed part of hypothesis 
ensures values likely to cause issues,
including small values, large values, zero, NaN, and Infinities
are handled appropriately.
In ensuring we handle all these different values,
Hypothesis is a tool that ensures that we codify
the assumptions we make about the data we put into our algorithms.&lt;&#x2F;p&gt;
&lt;p&gt;Working with hypothesis requires a fair amount of setup,
building upon a test framework like pytest.
Additionally, as a property testing framework,
rather than testing specific values
we instead need to test properties,
requires a rethinking of how we test functions.
In the example covered in the this article,
the property we are testing is ensuring our function
is doesn&#x27;t unexpectedly raise an exception.
The testing of this property
can be achieved by creating the test
as shown in the code snippet below.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;from hypothesis import given
from hypothesis.extra.numpy import arrays

@given(x=arrays(dtype=float, shape=10), y=arrays(dtype=float, shape=10))
def test_example(x, y):
    fit_linear_region(x, y)
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Here we use hypothesis to generate the two inputs to our function, &lt;code&gt;x&lt;&#x2F;code&gt; and &lt;code&gt;y&lt;&#x2F;code&gt;.
Each of these is created from a numpy array of 10 elements (&lt;code&gt;shape=10&lt;&#x2F;code&gt;)
containing floating point values (&lt;code&gt;dtype=float&lt;&#x2F;code&gt;).
The &lt;code&gt;@given&lt;&#x2F;code&gt; decorator allows hypothesis to generate inputs
based on the values that we have provided,
running each to check that it runs.
In this code snippet
we are making clear some assumptions that we have made about our data.
Firstly that we are expecting floating point values,
being the &lt;code&gt;dtype&lt;&#x2F;code&gt; that we are generating.
The second is that the &lt;code&gt;shape&lt;&#x2F;code&gt; of both &lt;code&gt;x&lt;&#x2F;code&gt; and &lt;code&gt;y&lt;&#x2F;code&gt; is equal.
One of the advantages of hypothesis
is bringing to light these occasionally hidden assumptions
about the inputs to our functions.&lt;&#x2F;p&gt;
&lt;p&gt;By using hypothesis to generate inputs to our function,
we find some additional errors that can occur.
When there are no points that lie within our specified range,
we get a &lt;code&gt;ValueError&lt;&#x2F;code&gt;,
while when there is only a single value in that range 
we get a &lt;code&gt;TypeError&lt;&#x2F;code&gt;.
Additionally, when scipy is unable to find
and appropriate line of best fit,
an &lt;code&gt;OptimizeWarning&lt;&#x2F;code&gt; is raised.
This results in the following &lt;code&gt;fit_linear_region&lt;&#x2F;code&gt; function
to handle these three exception cases.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;def fit_linear_region(x, y):
    &amp;quot;&amp;quot;&amp;quot;Fit a linear curve where the points lie within a region of 2D space.&amp;quot;&amp;quot;&amp;quot;
    # Select a region of points
    region = np.logical_and(2 &amp;lt; y, y &amp;lt; 50)
    # Perform the curve fitting only using the linear region
    try:
        popt, pcov = scipy.optimize.curve_fit(linear, x[region], y[region])
    except (ValueError, TypeError, scipy.optimize.OptimizeWarning):
        return np.nan, np.nan
    # Calculate the standard deviation of the fit parameters
    perr = 2 * np.sqrt(np.diag(pcov))
    # Returns the optimal curve and associated error
    return popt, perr
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;While using hypothesis allows us to easily find
some of the failure modes of functions,
it is another tool to learn how to use effectively.
It would be nicer if there was a simpler solution
to finding the errors raised by a function
that made use of existing tools.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;documentation&quot;&gt;Documentation&lt;&#x2F;h3&gt;
&lt;p&gt;The documentation for a function
is something regularly referred to,
whether that be through the online documentation
or through the help messages of a function.
This makes the function docstrings a fantastic place
to put information about the exceptions a function could raise.
The scipy docs include this information
for some of the functions, including &lt;a href=&quot;https:&#x2F;&#x2F;docs.scipy.org&#x2F;doc&#x2F;scipy&#x2F;reference&#x2F;generated&#x2F;scipy.optimize.curve_fit.html&quot;&gt;curve fit&lt;&#x2F;a&gt;.
Within the &lt;strong&gt;Raises&lt;&#x2F;strong&gt; section,
it notes there are three different exceptions the function can raise,
a &lt;code&gt;ValueError&lt;&#x2F;code&gt;, &lt;code&gt;RuntimeError&lt;&#x2F;code&gt;, or &lt;code&gt;OptimizeWarning&lt;&#x2F;code&gt;.
This list of exceptions demonstrates a limitation of using hypothesis,
it won&#x27;t always find &lt;em&gt;all&lt;&#x2F;em&gt; the possible exceptions a function can raise,
here missing the &lt;code&gt;RuntimeError&lt;&#x2F;code&gt;.
However, on the flip side, 
the &lt;code&gt;TypeError&lt;&#x2F;code&gt; we found with hypothesis is currently missing from the documentation.
This is a limitation of the documentation approach,
the exception comes from a function called by &lt;code&gt;curve_fit&lt;&#x2F;code&gt;,
propagating beyond the initial call site,
making it difficult to notice and to keep the documentation up to date.
The documentation approach is also prone to degrading over time,
as the code gets updated and changed,
the documentation has to be changed and updated alongside it,
and for exceptions this can be in multiple unrelated locations.
One of the solutions for this
is representing the possible exceptions in code,
rather than through the documentation
allowing tooling to provide assistance,
much like &lt;a href=&quot;http:&#x2F;&#x2F;mypy-lang.org&#x2F;&quot;&gt;mypy&lt;&#x2F;a&gt; and &lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;microsoft&#x2F;pyright&quot;&gt;pyright&lt;&#x2F;a&gt; do for type annotations.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;rust&quot;&gt;Rust&lt;&#x2F;h3&gt;
&lt;p&gt;One of the highly regarded parts of the Rust programming language,
is the comprehensive type system that explicitly handles errors.
In a similar way to exceptions within Python,
the &lt;code&gt;Result&lt;&#x2F;code&gt; type within Rust
is a way of stating that something has gone wrong.
The big difference between an exception in python
and the &lt;code&gt;Result&lt;&#x2F;code&gt; type in Rust,
is that in Rust you &lt;em&gt;must&lt;&#x2F;em&gt; explicitly handle the error.
In the simplest case this error handling is equivalent to a python exception,
stopping the execution of the program and exiting.
The key enhancement with the Rust approach
is that there is a record of every place
within the code that an error could halt the execution.
When prototyping on the sample datasets within Rust,
we can use the exception model of error handling,
halting the execution of the program using &lt;code&gt;.unwrap()&lt;&#x2F;code&gt;.
However, when moving to a more complex analysis pipeline,
or when developing a library that people rely on,
unexpected errors are no longer a good option.
Here, all instances of &lt;code&gt;.unwrap()&lt;&#x2F;code&gt;
can be converted into more appropriate methods
to handle those errors.
These propagation of the &lt;code&gt;Result&lt;&#x2F;code&gt; type occurs similarly to
the handling of &lt;code&gt;Optional&lt;&#x2F;code&gt; values within type checked python code.
Whenever we call a function that returns a &lt;code&gt;Result&lt;&#x2F;code&gt;,
we have to check whether we have the &lt;code&gt;Ok&lt;&#x2F;code&gt; value,
in which case we can continue on,
or the &lt;code&gt;Err&lt;&#x2F;code&gt; value,
where it has to be either be handled
or passed on to a calling function through the &lt;code&gt;Result&lt;&#x2F;code&gt; type.&lt;&#x2F;p&gt;
&lt;p&gt;In principle, supporting a &lt;code&gt;Result&lt;&#x2F;code&gt; type within python
like exists within Rust is possible, however,
what makes the type so useful within Rust
is that it &lt;em&gt;must&lt;&#x2F;em&gt; be handled,
a guarantee that can&#x27;t and won&#x27;t be enforced within python.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;conclusion&quot;&gt;Conclusion&lt;&#x2F;h2&gt;
&lt;p&gt;One of the key principles of python
is the ease of getting started.
Advanced parts of the language,
can be ignored until they are required.
Ignoring advanced parts of the language
is not just helpful for the learning process,
but also for slowly building up complexity
in the projects that we start.
When working with long running processes,
exception handling becomes an important part of the development process,
however, I don&#x27;t know of a clear way
to find all the possible ways an exception could be raised
from a function we might use.
This makes the development process a continually iterative one,
run the code until you find an exception,
handle the exception and start again,
a process that is not ideal,
particularly for overnight calculations.&lt;&#x2F;p&gt;
&lt;div class=&quot;footnote-definition&quot; id=&quot;1&quot;&gt;&lt;sup class=&quot;footnote-definition-label&quot;&gt;1&lt;&#x2F;sup&gt;
&lt;p&gt;It was only using 32 bit unsigned integers, however that was decided to be more than anyone would need.&lt;&#x2F;p&gt;
&lt;&#x2F;div&gt;
&lt;div class=&quot;footnote-definition&quot; id=&quot;2&quot;&gt;&lt;sup class=&quot;footnote-definition-label&quot;&gt;2&lt;&#x2F;sup&gt;
&lt;p&gt;If you have to know, I was calculating the Diffusion constant, 
a measure of how fast particles move over long timescales.
This is a linear function of the Mean Squared Displacement vs time.&lt;&#x2F;p&gt;
&lt;&#x2F;div&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>A Showcase of Data Analysis in Python and R: A Case Study using COVID-19 Data</title>
        <published>2020-08-11T00:00:00+00:00</published>
        <updated>2020-08-11T00:00:00+00:00</updated>
        <author>
          <name>Unknown</name>
        </author>
        <link rel="alternate" href="https://malramsay.com/talks/data-analysis-showcase/" type="text/html"/>
        <id>https://malramsay.com/talks/data-analysis-showcase/</id>
        
        <content type="html">&lt;div style=&quot;position: relative; padding-bottom: 56.25%; height: 0; overflow: hidden;&quot;&gt;
  &lt;iframe src=&quot;https:&#x2F;&#x2F;www.youtube.com&#x2F;embed&#x2F;DjHmaT0G770&quot;
    style=&quot;position: absolute; top: 0; left: 0; width: 100%; height: 100%; border: 0;&quot; title=&quot;youtube video&quot;
    webkitallowfullscreen mozallowfullscreen allowfullscreen&gt;
  &lt;&#x2F;iframe&gt;
&lt;&#x2F;div&gt;
&lt;p&gt;In all fields of research we are being confronted with a deluge of data; data that needs cleaning and transformation to be used in further analysis. This webinar demonstrates the effective use of programming tools for an initial analysis of COVID-19 datasets, with examples using both R and Python.&lt;&#x2F;p&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>Failing Hard: How I &#x27;lost&#x27; two years of data</title>
        <published>2018-11-22T00:00:00+00:00</published>
        <updated>2018-11-22T00:00:00+00:00</updated>
        <author>
          <name>Unknown</name>
        </author>
        <link rel="alternate" href="https://malramsay.com/post/failing-hard/" type="text/html"/>
        <id>https://malramsay.com/post/failing-hard/</id>
        
        <content type="html">&lt;p&gt;Failure is one of those topics
that is discussed far less than it should be
---It is hard to tell other people about mistakes you have made---
yet failure is usually a far better teacher than success.
This is why I want to share my story of
how I spent the first two years of my PhD collecting useless data
because of a bug in my code.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;the-problem&quot;&gt;The Problem&lt;&#x2F;h2&gt;
&lt;p&gt;I perform computer simulations of toy molecular systems,
using something simple and understandable
to gain insights into more complicated chemical systems.
This is just like in High School and Undergraduate Physics
where you perform calculations for objects in a frictionless vacuum.
Apart from the simplicity of a toy system,
a major benefit of using one is that you can
modify individual parameters of the system to test a hypothesis.
I had the hypothesis that rotational motion plays an important role
in the large scale translational motion---known as &lt;em&gt;diffusion&lt;&#x2F;em&gt;,
so I was using a toy system which allowed me to control
the effective magnitudes of these two types of motion.
The small scale translational motion is related to the &lt;em&gt;mass&lt;&#x2F;em&gt;;
heavy objects are more difficult to move in a straight line.
While the rotational motion is related to the &lt;em&gt;moment-of-inertia&lt;&#x2F;em&gt;,
a measure of how difficult something is to rotate.
Though varying these two properties,
the &lt;em&gt;mass&lt;&#x2F;em&gt; and the &lt;em&gt;moment-of-inertia&lt;&#x2F;em&gt;,
I can establish how the translational and rotational motions
contribute to the rate of diffusion.
Setting the &lt;em&gt;moment-of-inertia&lt;&#x2F;em&gt; really large
effectively stops the rotational motion,
so what happens to the diffusion?
Or the opposite, making the &lt;em&gt;moment-of-inertia&lt;&#x2F;em&gt; really small
what is the resulting effect on the diffusion?&lt;&#x2F;p&gt;
&lt;p&gt;Before I start changing the parameters of the toy model,
I need a reference (or control) model related to existing results.
Science is always performed with reference to our current understanding,
so the control model provides a point of comparison to existing results,
including other toy models and real chemical systems.
The molecule I am studying is comprised of three circles,
one large and two small, arranged like a Mickey Mouse head.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;img src=&quot;&#x2F;img&#x2F;molecules&#x2F;trimer.png&quot; alt=&quot;The molecule I am studying, three circles arranged like a Mickey Mouse head.&quot; &#x2F;&gt;&lt;&#x2F;p&gt;
&lt;p&gt;To make calculations with these toy systems easier,
I use a system of units in which most quantities have a unitless value of 1.
Each particle has a point mass at the center of the circle of 1,
the large circle has a radius of 1,
a reasonable temperature is 1
---It is really nice that most of the multiplication and division disappears.
The smaller particles in the molecule have a radius of $0.63$
and are located on the circumference of the large particle
at an angle of $120^\circ$.&lt;&#x2F;p&gt;
&lt;p&gt;I am using the above molecule in Molecular Dynamics (MD) simulations,
which use Newton&#x27;s equations of motion to update the positions of each particle.
The translational acceleration $a$ of a molecule
is related to the force $F$ acting on it divided by the mass $m$
as given by the equation $a=F&#x2F;m$.
While the angular acceleration $\alpha$
is related to the torque $\tau$ divided by the angular momentum $I$
as given by $\alpha = \tau&#x2F;I$.
Molecular Dynamics simulations have two different methods
of dealing with the molecules I am using.
The first is to solve these equations for
each of a molecule&#x27;s component particles individually,
followed by a second step which restores the molecular shape.
The second approach is
to calculate the force on each component particle,
then adding them together
to solve the above equations for the molecule as a whole.&lt;&#x2F;p&gt;
&lt;p&gt;Molecular Dynamics simulations are a standard computational tool
for understanding a variety of problems
from crystal growth to protein folding.
As a widely used tool,
there are many different software packages available
for performing these simulations,
with each having their own strengths and weaknesses.
Using these software packages
involves expressing the simulation you want to perform
in the appropriate manner for the software.
At the start of my PhD
I decided that &lt;a href=&quot;https:&#x2F;&#x2F;hoomd-blue.readthedocs.io&#x2F;en&#x2F;stable&#x2F;index.html&quot;&gt;Hoomd&lt;&#x2F;a&gt; developed by the Glotzer group
at the University of Michigan was the most suitable
for the types of simulations I was going to be performing.
Soon after starting to understand the software,
there was a new major release,
which included a significant overhaul of the software.
This overhaul included changing how it handled the molecular calculations
from the first method to the second method,
something which I didn&#x27;t realise at the time.
In getting my simulations to work with the new version of Hoomd,
I manually entered a moment of inertia,
calculated from each particle having a point mass of 1,
however the total mass of the molecule
was taken as the mass of an individual particle being 1 instead of 3.
By not actually checking the simulations were behaving as I intended
I spent the next two years of my PhD
characterising the behaviour of the wrong control experiment.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;finding-the-bug&quot;&gt;Finding the Bug&lt;&#x2F;h2&gt;
&lt;p&gt;I found the bug as part of trying to add additional functionality.
I was trying to randomly initialise velocities of particles in a configuration
such that they matched the desired temperature.&lt;&#x2F;p&gt;
&lt;p&gt;The reason I found the bug was a result few the
Bug result of setting thermodynamic quantities
- temperature is a distribution of velocities
- energy is evenly distributed between translational and rotational motion&lt;&#x2F;p&gt;
&lt;h2 id=&quot;lessons-learnt&quot;&gt;Lessons Learnt&lt;&#x2F;h2&gt;
&lt;p&gt;The lessons I have learnt from this experience
can be grouped into two groups,
the identification of bugs,
and the management of computational projects.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;identifying-bugs&quot;&gt;Identifying Bugs&lt;&#x2F;h3&gt;
&lt;p&gt;Even before the identification of this bug
I had developed a test suite for my code,
ensuring I was configuring the simulations in Hoomd
with the appropriate values.
What I wasn&#x27;t checking,
was that these values produced the simulations I thought I was running.
Typically the method of doing this
is to take some previous results and replicating them,
although since no-one else has run simulations
on the particular system I am using it is a little difficult to do this.&lt;&#x2F;p&gt;
&lt;p&gt;I had been told by my supervisor on many occasions
to confirm that the simulations were physically accurate,
usually with reference to conservation of energy.
This was really easy to ignore because that out of scope
for the code I was writing,
however, the idea of checking the simulation
is consistent with the laws of physics
was definitely worth considering earlier.&lt;&#x2F;p&gt;
&lt;blockquote&gt;
&lt;p&gt;Testing is not about ensuring exact correctness,
it is also ensuring the absence of obvious errors.&lt;&#x2F;p&gt;
&lt;&#x2F;blockquote&gt;
&lt;p&gt;A simple check of physicality is though the equi partition theorem,
which describes that each of the degrees of freedom,
whether translational or rotational,
have the same energy.
The relationships below link the mass $m$
and moment-of-inertia $I$
to the temperature $T$ providing a route to testing the simulations,
where the angled brackets $\langle \rangle$ denote an averaged quantity.&lt;&#x2F;p&gt;
&lt;p&gt;\[
\begin{aligned}
\langle \frac{1}{2} m v^2 \rangle &amp;amp;= k_B T \\
\langle \frac{1}{2} I \omega^2 \rangle &amp;amp;= \frac{1}{2}k_B T
\end{aligned}
\]&lt;&#x2F;p&gt;
&lt;p&gt;What might not be so obvious here is that the temperature $T$
in the above equations is not a static quantity,
instead fluctuating about the desired temperature,
a construct of the type of simulations I am performing.
This makes the task of testing the temperature is correct
significantly more difficult,
since it is not a set value,
lying in a distribution.
While creating automated testing for whether a value that varies can be difficult,
re-framing the test as &lt;em&gt;not-wrong&lt;&#x2F;em&gt; is just as valid.
With this bug I created I was off by a factor of 3,
many times larger than any reasonable variation.
I didn&#x27;t need to test that the temperature was exactly right,
I just needed to ensure it is not obviously wrong.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;rebuilding-the-dataset&quot;&gt;Rebuilding the Dataset&lt;&#x2F;h3&gt;
&lt;p&gt;While the main storyline of this article has been
losing the years of my PhD research,
a major part of the learning experience
has been the rebuilding of all the data I had collected.&lt;&#x2F;p&gt;
&lt;blockquote&gt;
&lt;p&gt;Reproducibility in science is not just for others to reproduce your work
it is also so you can reproduce it when you stuff up.&lt;&#x2F;p&gt;
&lt;&#x2F;blockquote&gt;
&lt;p&gt;I have been constantly working to improve my workflow,
as I learn more about what works and what doesn&#x27;t
ending up with a structure based heavily on
the &lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;MolSSI&#x2F;cookiecutter-cms&quot;&gt;cookiecutter-cms&lt;&#x2F;a&gt; and &lt;a href=&quot;https:&#x2F;&#x2F;drivendata.github.io&#x2F;cookiecutter-data-science&#x2F;&quot;&gt;cookiecutter-data-science&lt;&#x2F;a&gt; project templates.
All the new work was structured into these project templates
which are well organised, with an emphasis on being able to replicate analyses.
What I really noticed in rebuilding my dataset
was that replicating these newer elements was simple and straightforward,
while the older projects just working out what I had done was a challenge,
let alone what needed to be recreated.&lt;&#x2F;p&gt;
&lt;p&gt;Further to having a well organised data analysis,
having all the computational experiments,
including the permutations of every variable
and the scripts required to run the simulation was invaluable.
The experiments where this was up to date
were the simplest to deal with,
I just needed to run them again.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;conclusion&quot;&gt;Conclusion&lt;&#x2F;h2&gt;
&lt;p&gt;Although we often don&#x27;t like to admit it, we all make mistakes.
Mistakes, failures, and bugs are all part of doing research.
While they may be disruptive and inconvenient,
they are only truly damaging when they keep re-occurring.
Making a mistake once is part of the process,
repeatedly making the same mistake is a problem.
Best practices like automated testing
are ways of making the mistakes that do occur obvious and simple to fix,
as well as providing a way of preventing them from reoccurring.
I am now fairly confident that should this particular bug reappear I will notice,
although that doesn&#x27;t mean there isn&#x27;t another hiding away for me to find.&lt;&#x2F;p&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>A Complete Guide to Docker on Fedora</title>
        <published>2018-09-02T00:00:00+00:00</published>
        <updated>2018-09-02T00:00:00+00:00</updated>
        <author>
          <name>Unknown</name>
        </author>
        <link rel="alternate" href="https://malramsay.com/post/docker-on-fedora/" type="text/html"/>
        <id>https://malramsay.com/post/docker-on-fedora/</id>
        
        <content type="html">&lt;p&gt;While there are numerous guides to installing Docker on Fedora,
none of the guides leave the installation
in a state that I would consider usable.
This is intended to be a single complete guide
for the setup and configuration of Docker,
highlighting the differences that are required
to get Docker running on Fedora.
I will be demonstrating using Fedora 28,
however, this should be the same for previous or future releases.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;installation&quot;&gt;Installation&lt;&#x2F;h2&gt;
&lt;p&gt;This is the part that all the guides include,
including the &lt;a href=&quot;https:&#x2F;&#x2F;developer.fedoraproject.org&#x2F;tools&#x2F;docker&#x2F;docker-installation.html&quot;&gt;fedora documentation&lt;&#x2F;a&gt;.
Docker is in the Fedora repositories
enabling installation using the &lt;code&gt;dnf&lt;&#x2F;code&gt; package manager&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;sudo dnf install docker
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Once installed, the Docker service can be started by running&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;sudo systemctl start docker
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;and should you want to start docker
every time you boot your machine
you can run&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;sudo systemctl enable docker
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Note that the above command doesn&#x27;t start the Docker service immediately,
so you will have to run both the &lt;code&gt;start&lt;&#x2F;code&gt; and &lt;code&gt;enable&lt;&#x2F;code&gt; commands
to have the Docker service running now and on following reboots.&lt;&#x2F;p&gt;
&lt;p&gt;At this point you might want to try running a Docker container&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$ docker run hello-world
&amp;#x2F;usr&amp;#x2F;bin&amp;#x2F;docker-current: Got permission denied while trying to connect
to the Docker daemon socket at unix:&amp;#x2F;&amp;#x2F;&amp;#x2F;var&amp;#x2F;run&amp;#x2F;docker.sock ...
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;only you get a permission denied error.
It is possible to run Docker as &lt;code&gt;root&lt;&#x2F;code&gt;,
however it is probably not the best idea
since it is kind of simple to make a mistake.&lt;&#x2F;p&gt;
&lt;p&gt;If instead you received the message&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$ docker run hello-world
&amp;#x2F;usr&amp;#x2F;bin&amp;#x2F;docker-current: Cannot connect to the Docker daemon at
unix:&amp;#x2F;&amp;#x2F;&amp;#x2F;var&amp;#x2F;run&amp;#x2F;docker.sock. Is the docker daemon running?
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;this means you haven&#x27;t started Docker and need to run&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;sudo systemctl start docker
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;h2 id=&quot;setting-permissions&quot;&gt;Setting Permissions&lt;&#x2F;h2&gt;
&lt;blockquote class=&quot;warning&quot;&gt;
    &lt;p&gt;This section has the potential to break things which are difficult to fix.
Please be really careful, unlike me.&lt;&#x2F;p&gt;

&lt;&#x2F;blockquote&gt;
&lt;p&gt;This follows the optional &lt;a href=&quot;https:&#x2F;&#x2F;docs.docker.com&#x2F;install&#x2F;linux&#x2F;linux-postinstall&#x2F;&quot;&gt;post installation&lt;&#x2F;a&gt; steps in the Docker documentation.&lt;&#x2F;p&gt;
&lt;p&gt;For users to have permission to access the Docker socket,
they either need to be root,
or they can be a member of the &lt;code&gt;docker&lt;&#x2F;code&gt; group.&lt;&#x2F;p&gt;
&lt;p&gt;This group probably doesn&#x27;t exist on your system yet,
though you can check by running&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;grep docker &amp;#x2F;etc&amp;#x2F;group
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;If there is no output the group does not yet exist
and can be created with the &lt;code&gt;groupadd&lt;&#x2F;code&gt; command&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;sudo groupadd docker
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;The group should now appear in the &lt;code&gt;&#x2F;etc&#x2F;group&lt;&#x2F;code&gt; file&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$ grep docker &amp;#x2F;etc&amp;#x2F;group
docker:x:1001:
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;The final step is adding yourself and&#x2F;or any other users to the Docker group.
This is done with the command&lt;&#x2F;p&gt;
&lt;blockquote class=&quot;warning&quot;&gt;
    &lt;p&gt;Running the below command without append
will remove you from the &lt;code&gt;wheel&lt;&#x2F;code&gt; group
meaning you will no longer be able to run commands with &lt;code&gt;sudo&lt;&#x2F;code&gt;.
If you are the only user with root access
you will have to repair your install from a live image.&lt;&#x2F;p&gt;

&lt;&#x2F;blockquote&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;sudo usermod --append --groups docker $USER
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;You have to go through the login process to update your group membership,
with the safest method being to open an &lt;code&gt;ssh&lt;&#x2F;code&gt; connection to &lt;code&gt;localhost&lt;&#x2F;code&gt;.
This means that if you accidentally remove yourself from the &lt;code&gt;wheel&lt;&#x2F;code&gt; group,&lt;sup class=&quot;footnote-reference&quot;&gt;&lt;a href=&quot;#1&quot;&gt;1&lt;&#x2F;a&gt;&lt;&#x2F;sup&gt;
you just have to disconnect the session
to regain sudo permissions and fix things.&lt;&#x2F;p&gt;
&lt;p&gt;To check everything is as expected,
the &lt;code&gt;groups&lt;&#x2F;code&gt; command will list the groups you are a part of.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$ groups
malcolm wheel docker
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;You should have a list of groups similar to those output above.
Now you are a member of the &lt;code&gt;docker&lt;&#x2F;code&gt; group
you can test Docker is working with the test image&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$ docker run hello-world

Hello from Docker!
This message shows that your installation appears to be working correctly.

...
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;This indicates that Docker is working properly.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;mounting-local-directories&quot;&gt;Mounting Local Directories&lt;&#x2F;h2&gt;
&lt;p&gt;Now you have Docker working,
you probably want to do something useful with it
like access the local filesystem for processing.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$ docker run --interactive --tty --volume $(pwd):&amp;#x2F;srv ubuntu
root@aa7c5f9dcef9:&amp;#x2F;# _
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;The above command creates an interactive terminal (tty)
running in an Ubuntu container,
with the prompt for the container now showing.
Additionally we have mounted the current directory to the container
at the &lt;code&gt;&#x2F;srv&lt;&#x2F;code&gt; folder of the container.
We can try and access the contents of the current directory
from within the container&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;root@aa7c5f9dcef9:&amp;#x2F;# ls
ls: can&amp;#x27;t open &amp;#x27;&amp;#x2F;srv&amp;#x27;: Permission denied
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;On an Ubuntu install this would work,
however Fedora uses SELinux for security,
which requires the appropriate labelling of file objects
for the processes using them.
By default Docker doesn&#x27;t perform this labelling,
however we can tell it to with the &lt;code&gt;:z&lt;&#x2F;code&gt; or &lt;code&gt;:Z&lt;&#x2F;code&gt; suffixes for the volume.
The lowercase &lt;code&gt;:z&lt;&#x2F;code&gt; allows multiple containers to access the volume
and the uppercase &lt;code&gt;:Z&lt;&#x2F;code&gt; allows a single container to access the volume.
The command becomes&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$ docker run -it -v $(pwd):&amp;#x2F;srv:Z ubuntu
root@aa7c5f9dcef9:&amp;#x2F;# ls &amp;#x2F;srv
docker_on_fedora.md
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Here I have used the more common shortened command line options,
&lt;code&gt;-it&lt;&#x2F;code&gt; for the interactive terminal, and &lt;code&gt;-v&lt;&#x2F;code&gt; for the volume.&lt;&#x2F;p&gt;
&lt;p&gt;For a program that at first glance appears to simple to install,
Docker is rather difficult to get set up properly on Fedora.
Hopefully this&lt;&#x2F;p&gt;
&lt;div class=&quot;footnote-definition&quot; id=&quot;1&quot;&gt;&lt;sup class=&quot;footnote-definition-label&quot;&gt;1&lt;&#x2F;sup&gt;
&lt;p&gt;I have done this...twice&lt;&#x2F;p&gt;
&lt;&#x2F;div&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>Speed Through Specificity</title>
        <published>2018-08-07T00:00:00+00:00</published>
        <updated>2018-08-07T00:00:00+00:00</updated>
        <author>
          <name>Unknown</name>
        </author>
        <link rel="alternate" href="https://malramsay.com/post/speed-through-specificity/" type="text/html"/>
        <id>https://malramsay.com/post/speed-through-specificity/</id>
        
        <content type="html">&lt;p&gt;A common criticism about the Python programming language is that it is slow, often with reference to
a benchmark comparing a range of tasks. This criticism is widely addressed with articles by
&lt;a href=&quot;https:&#x2F;&#x2F;jakevdp.github.io&#x2F;blog&#x2F;2014&#x2F;05&#x2F;09&#x2F;why-python-is-slow&#x2F;&quot;&gt;Jake van der Plass&lt;&#x2F;a&gt; and &lt;a href=&quot;https:&#x2F;&#x2F;hackernoon.com&#x2F;why-is-python-so-slow-e5074b6fe55b&quot;&gt;Anthony Shaw&lt;&#x2F;a&gt; being two excellent examples. While I don&#x27;t disagree with any of
the points raised in these articles, I think they miss an important aspect of
performance---specificity. Python is a general purpose language, used for nearly everything from
embedded devices with &lt;a href=&quot;https:&#x2F;&#x2F;micropython.org&#x2F;&quot;&gt;uPython&lt;&#x2F;a&gt; to distributed processing of &lt;a href=&quot;https:&#x2F;&#x2F;www.youtube.com&#x2F;watch?v=Hd_ydJeyr5M&quot;&gt;petabytes of data&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;Speed comes through specificity for a task. Even when working with C or C++, which are generally
regarded as the reference standard for performance, there is still an argument to be made for
increased performance through hand optimised assembly. Writing assembly, which is the stream of
instructions the CPU interprets to do it&#x27;s work, &lt;em&gt;can&lt;&#x2F;em&gt; result in faster code than a compiler if you
know what you doing. Only nearly no-one actually writes assembly because we want our applications to
work on different processor architectures and to take advantage of the features of newer CPUs like
out-of-order execution.&lt;sup class=&quot;footnote-reference&quot;&gt;&lt;a href=&quot;#1&quot;&gt;1&lt;&#x2F;a&gt;&lt;&#x2F;sup&gt; Instead of writing the fastest possible code for a particular processor
a more common approach is to provide hints to the compiler to help it optimise performance, trading
some performance for generality and simplicity.&lt;&#x2F;p&gt;
&lt;p&gt;While it is uncommon to try and eke out every last bit of performance from a CPU, offloading the
work to more specialised hardware is commonplace. A CPU is a general purpose tool being adaptable to
many different types of operations which make it ideal for powering our computers, though to really
understand &lt;em&gt;fast&lt;&#x2F;em&gt; we need specialised hardware like GPUs, FPGAs, or ASICs. Graphics Processing Units
(&lt;a href=&quot;https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Graphics_processing_unit&quot;&gt;GPUs&lt;&#x2F;a&gt;)&lt;sup class=&quot;footnote-reference&quot;&gt;&lt;a href=&quot;#2&quot;&gt;2&lt;&#x2F;a&gt;&lt;&#x2F;sup&gt; are designed for performing the same operation on large quantities of data at the same
time; whether that is working out the colour of each pixel on a display, or the values at each point
of a large matrix. GPUs are designed at a hardware level to perform these types of tasks, eschewing
much of the capability of a CPU. Having hardware specific to a task is a key factor for performance
and Field Programmable Gate Arrays (&lt;a href=&quot;https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Field-programmable_gate_array&quot;&gt;FPGAs&lt;&#x2F;a&gt;) are one method of achieving this. Instead of providing
a stream of instructions to be interpreted like a CPU or GPU, an FPGA is programmed by arranging the
circuits to perform the desired processing, providing phenomenal processing capability. FPGAs are
used in places like &lt;a href=&quot;https:&#x2F;&#x2F;people.eecs.berkeley.edu&#x2F;~bora&#x2F;publications&#x2F;Asilomar06b.pdf&quot;&gt;signal processing&lt;&#x2F;a&gt; and the &lt;a href=&quot;https:&#x2F;&#x2F;www.eetimes.com&#x2F;document.asp?doc_id=1262350&quot;&gt;Mars Rovers&lt;&#x2F;a&gt;
as they allow for reprogramming for task switching or hardware updates. In some applications
the programmability of a FPGA is unnecessary, so Application-Specific Integrated Circuits (&lt;a href=&quot;https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Application-specific_integrated_circuit&quot;&gt;ASICs&lt;&#x2F;a&gt;)
are used instead. An ASIC is a piece of silicon for processing a single specific task, with a common
use case being decoding video streams enabling you to watch hours of video in a single charge on
your phone. The progression of hardware from CPU to GPU to FPGA to ASIC represents an increasing
specificity to a task for improved performance at that task.&lt;&#x2F;p&gt;
&lt;p&gt;The specificity of a task doesn&#x27;t just refer to the low level details of processor architecture,
there are also the levels most developers are more likely to encounter of type and application
specificity. Python is a dynamically typed programing language, allowing variables to change type
and change functionality during execution. A byproduct of this dynamic typing is types and operator
functions need to be evaluated for each operation. Should we want to add the values of two lists
together like in the example below, each time the code reaches the line &lt;code&gt;result = i + j&lt;&#x2F;code&gt; it has to
evaluate the types of both &lt;code&gt;i&lt;&#x2F;code&gt; and &lt;code&gt;j&lt;&#x2F;code&gt; to then work out how to perform the addition operation. Most
noticeable with large numbers of numerical values, the type evaluation takes far longer than the
addition operation.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;list1 = [1, 2, 3]
list2 = [4, 5, 6]
list_result = []
for i, j in zip(list1, list2):
    result = i + j
    list_result.append(result)
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;A key concept in optimising numerical python code is limiting the number of type evaluations, with
the canonical method being to use &lt;a href=&quot;http:&#x2F;&#x2F;www.numpy.org&#x2F;&quot;&gt;NumPy&lt;&#x2F;a&gt; arrays. Instead of having a list which is able to hold
many different objects, NumPy arrays only contain a single type meaning the type only needs to be
evaluated once for the entire array. Other optimisation techniques for numerical python, namely
&lt;a href=&quot;http:&#x2F;&#x2F;cython.org&#x2F;&quot;&gt;Cython&lt;&#x2F;a&gt; and &lt;a href=&quot;https:&#x2F;&#x2F;numba.pydata.org&#x2F;&quot;&gt;Numba&lt;&#x2F;a&gt;, operate in a very similar way; reducing the calculation to a limited range of
types evaluated once. With these tools for numerical computation it is possible to get performance
on par or even exceeding that of C code.&lt;&#x2F;p&gt;
&lt;p&gt;Even with the most specific type information the approach to solving a problem is a huge factor for
performance. Applications designed to solve a very specific problem are able to make significant
optimisations though handling only the cases the application will come across; an algorithm for
acyclic graphs doesn&#x27;t have to handle cyclic graphs. Additionally an application for user
interaction can optimise for responsiveness, or a data processing application can optimise for data
throughput. The general purpose tool has to optimise for everything, so it is optimised for nothing.
Python is a general purpose programming language; not optimised for numerical computing, for the
web, for command line scripts, or any other use case. However, the adaptability of python has
allowed the development of application specific tools, like NumPy, containing a small range of very
specific functionality. Python also allows for the simple integration of problem specific hardware,
with tools like &lt;a href=&quot;https:&#x2F;&#x2F;documen.tician.de&#x2F;pycuda&#x2F;&quot;&gt;PyCUDA&lt;&#x2F;a&gt;, &lt;a href=&quot;https:&#x2F;&#x2F;scikit-cuda.readthedocs.io&#x2F;en&#x2F;latest&#x2F;&quot;&gt;scikit-cuda&lt;&#x2F;a&gt;, &lt;a href=&quot;http:&#x2F;&#x2F;www.myhdl.org&#x2F;&quot;&gt;MyHDL&lt;&#x2F;a&gt;, &lt;a href=&quot;http:&#x2F;&#x2F;nifpga-python.readthedocs.io&#x2F;en&#x2F;latest&#x2F;&quot;&gt;nifpga-python&lt;&#x2F;a&gt;, and many others.&lt;&#x2F;p&gt;
&lt;p&gt;Where speed is important, use the specific tool for the job. In most cases this isn&#x27;t the python
standard library, instead it is usually just a &lt;code&gt;pip install&lt;&#x2F;code&gt; away.&lt;sup class=&quot;footnote-reference&quot;&gt;&lt;a href=&quot;#3&quot;&gt;3&lt;&#x2F;a&gt;&lt;&#x2F;sup&gt;&lt;&#x2F;p&gt;
&lt;div class=&quot;footnote-definition&quot; id=&quot;1&quot;&gt;&lt;sup class=&quot;footnote-definition-label&quot;&gt;1&lt;&#x2F;sup&gt;
&lt;p&gt;I guess you could also consider this a &lt;a href=&quot;https:&#x2F;&#x2F;meltdownattack.com&#x2F;&quot;&gt;bug...&lt;&#x2F;a&gt;&lt;&#x2F;p&gt;
&lt;&#x2F;div&gt;
&lt;div class=&quot;footnote-definition&quot; id=&quot;2&quot;&gt;&lt;sup class=&quot;footnote-definition-label&quot;&gt;2&lt;&#x2F;sup&gt;
&lt;p&gt;I would consider &lt;a href=&quot;https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Tensor_processing_unit&quot;&gt;Tensor Processing Units&lt;&#x2F;a&gt; (TPUs) in the same category as GPUs, with the main difference being the targeted precision of mathematical operations.&lt;&#x2F;p&gt;
&lt;&#x2F;div&gt;
&lt;div class=&quot;footnote-definition&quot; id=&quot;3&quot;&gt;&lt;sup class=&quot;footnote-definition-label&quot;&gt;3&lt;&#x2F;sup&gt;
&lt;p&gt;Or &lt;code&gt;conda install&lt;&#x2F;code&gt; for the tools requiring C&#x2F;C++ libraries&lt;&#x2F;p&gt;
&lt;&#x2F;div&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>Compiling LaTeX on Travis-CI</title>
        <published>2018-07-16T00:00:00+00:00</published>
        <updated>2018-07-16T00:00:00+00:00</updated>
        <author>
          <name>Unknown</name>
        </author>
        <link rel="alternate" href="https://malramsay.com/post/compiling-latex-on-travis/" type="text/html"/>
        <id>https://malramsay.com/post/compiling-latex-on-travis/</id>
        
        <content type="html">&lt;p&gt;One of the best parts of the current software development environment is the proliferation of
Continuous Integration (CI) services like &lt;a href=&quot;https:&#x2F;&#x2F;travis.org&quot;&gt;Travis-CI&lt;&#x2F;a&gt;. These CI services plug into GitHub or other
code repositories to automatically run when new code is pushed to a repository. Typically CI is
used for running automated testing every time new code is added so you can be reasonably confident
a change hasn&#x27;t broken any functionality. The premise of CI is the automation of tedious tasks like
running tests.&lt;&#x2F;p&gt;
&lt;p&gt;When writing a LaTeX document, I find compilation the most tedious task. Particularly for large
documents, where it can take a long time. This means I run the compilation irregularly, invariably
resulting in a whole collection of errors that have accumulated and I now have to fix. Additionally
the compilation of LaTeX documents is often highly machine dependent, only working properly with a
specific configuration some reason. My final issue with LaTeX documents is ensuring all the files
required for compilation are included in the git repository, not just hiding in a directory
somewhere on your local filesystem.&lt;&#x2F;p&gt;
&lt;p&gt;As a way of preventing these issues in the write-up of my PhD thesis, I have developed a
configuration for building LaTeX documents on Travis-CI. While I have come across other methods for
compiling LaTeX documents on Travis-CI, they all compromise on the workflow I would like. I want
a system that is;&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;em&gt;fast&lt;&#x2F;em&gt;, with builds completing in a couple of minutes&lt;&#x2F;li&gt;
&lt;li&gt;&lt;em&gt;adaptable&lt;&#x2F;em&gt;, not having to manually specify every package I use&lt;&#x2F;li&gt;
&lt;li&gt;&lt;em&gt;extensible&lt;&#x2F;em&gt;, I can easily use &lt;a href=&quot;https:&#x2F;&#x2F;pandoc.org&#x2F;&quot;&gt;pandoc&lt;&#x2F;a&gt; to convert files to LaTeX before compilation&lt;&#x2F;li&gt;
&lt;li&gt;uses &lt;code&gt;biber&lt;&#x2F;code&gt;, the best practice for compiling the references&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;h2 id=&quot;choosing-a-latex-distribution&quot;&gt;Choosing a LaTeX Distribution&lt;&#x2F;h2&gt;
&lt;p&gt;There are a number of different methods to get a distribution of LaTeX installed. The main method
for Linux is the TeXLive distribution which is typically installed via the package manager. The
base TeXLive distribution is pre-installed on the Travis-CI image for A. However, this is only the
core extensions and so this approach is either &lt;em&gt;fast&lt;&#x2F;em&gt; in that the image just boots up, or
&lt;em&gt;adaptable&lt;&#x2F;em&gt; by downloading &lt;code&gt;texlive-full&lt;&#x2F;code&gt; which takes a long time.&lt;&#x2F;p&gt;
&lt;p&gt;The key issue with the TeXLive distribution is the lack of automatically downloading required
packages. A LaTeX distribution with this feature is &lt;a href=&quot;https:&#x2F;&#x2F;miktex.org&#x2F;&quot;&gt;MiKTeX&lt;&#x2F;a&gt;, having an installer that is only 200
MB compared to the ~3 GB of the complete TeXLive. MiKTeX also provides a
&lt;a href=&quot;https:&#x2F;&#x2F;miktex.org&#x2F;howto&#x2F;miktex-docker&quot;&gt;docker container&lt;&#x2F;a&gt; which is a great method for having exactly the same environment
compiling locally and through a CI service. Unfortunately MiKTeX doesn&#x27;t install the &lt;code&gt;biber&lt;&#x2F;code&gt; binary
on Linux (or macOS).&lt;sup class=&quot;footnote-reference&quot;&gt;&lt;a href=&quot;#1&quot;&gt;1&lt;&#x2F;a&gt;&lt;&#x2F;sup&gt; While it is possible to create a new container which includes the &lt;code&gt;biber&lt;&#x2F;code&gt;
binary, extending and maintaining a Docker container is non-trivial making this approach less
favourable. That said, for other Docker-centric CI services this could be an excellent approach.&lt;&#x2F;p&gt;
&lt;p&gt;Another newer and less well known LaTeX distribution is &lt;a href=&quot;https:&#x2F;&#x2F;tectonic-typesetting.github.io&quot;&gt;Tectonic&lt;&#x2F;a&gt;. Although still considered beta
software, it works for most scenarios and has a lot of features that make it suitable for CI. I
would recommend installation with &lt;a href=&quot;https:&#x2F;&#x2F;tectonic-typesetting.github.io&#x2F;en-US&#x2F;install.html#the-anaconda-method&quot;&gt;conda&lt;&#x2F;a&gt; although there are a range of
&lt;a href=&quot;https:&#x2F;&#x2F;tectonic-typesetting.github.io&#x2F;en-US&#x2F;install.html&quot;&gt;installation methods&lt;&#x2F;a&gt; for both Linux and macOS (currently no Windows support).
Like MiKTeX, Tectonic automatically downloads the packages required to compile a document, making
the same configuration adaptable to many different documents. Having conda as an installation method
is also particularly useful, allowing simple installation of many other tools (like &lt;a href=&quot;https:&#x2F;&#x2F;pandoc.org&#x2F;&quot;&gt;pandoc&lt;&#x2F;a&gt;) which
might be required to compile a more complex document. The only requirement not satisfied by Tectonic
is the automatic installation of &lt;code&gt;biber&lt;&#x2F;code&gt;. Although not supported natively, it is &lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;tectonic-typesetting&#x2F;tectonic&#x2F;issues&#x2F;53&quot;&gt;possible to use
biber with tectonic&lt;&#x2F;a&gt; as long as the appropriate binary for biber 2.5 is installed,
either from &lt;a href=&quot;https:&#x2F;&#x2F;sourceforge.net&#x2F;projects&#x2F;biblatex-biber&#x2F;files&#x2F;biblatex-biber&#x2F;2.5&#x2F;binaries&#x2F;&quot;&gt;Sourceforge&lt;&#x2F;a&gt; or using conda&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;bash&quot; class=&quot;language-bash &quot;&gt;&lt;code class=&quot;language-bash&quot; data-lang=&quot;bash&quot;&gt;conda install -c malramsay biber==2.5
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;h3 id=&quot;compiling-documents-with-tectonic&quot;&gt;Compiling Documents with Tectonic&lt;&#x2F;h3&gt;
&lt;p&gt;Typically, compiling documents with &lt;code&gt;tectonic&lt;&#x2F;code&gt; requires a single command&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;bash&quot; class=&quot;language-bash &quot;&gt;&lt;code class=&quot;language-bash&quot; data-lang=&quot;bash&quot;&gt;tectonic document.tex
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;creating the file &lt;code&gt;document.pdf&lt;&#x2F;code&gt; and  automatically removing all intermediate files normally
associated with compiling LaTeX documents. To use &lt;code&gt;biber&lt;&#x2F;code&gt; instead of &lt;code&gt;biblatex&lt;&#x2F;code&gt; for the bibliography
the process is not quite so simple. In an attempt to make it easier I have created the Makefile
below which can be used to create a document.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;Makefile&quot; class=&quot;language-Makefile &quot;&gt;&lt;code class=&quot;language-Makefile&quot; data-lang=&quot;Makefile&quot;&gt;# Makefile

# directory to put build files
build_dir := output

.PHONY: all

all: document.pdf

%.pdf: %.tex | $(build_dir)
	tectonic -o $(build_dir) --keep-intermediates -r0 $&amp;lt;
	# Run biber if we find a .bcf file in the output
	if [ -f $(build_dir)&amp;#x2F;$(notdir $(&amp;lt;:.tex=.bcf)) ]; then \
		biber --input-directory $(build_dir) $(notdir $(&amp;lt;:.tex=)); \
	fi
	tectonic -o $(build_dir) --keep-intermediates $&amp;lt;
	cp $(build_dir)&amp;#x2F;$(notdir $@) .

$(build_dir):
	mkdir -p $@
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;This build process runs &lt;code&gt;tectonic&lt;&#x2F;code&gt; once with the &lt;code&gt;--keep-intermediates&lt;&#x2F;code&gt; option to generate the
intermediate files. I then check for the presence of a &lt;code&gt;.bcf&lt;&#x2F;code&gt; file, which is the file &lt;code&gt;biber&lt;&#x2F;code&gt; uses
to to it&#x27;s thing. Tectonic is then run afterwards, which runs the compilation step as many times as
it needs to finalise the output. The final step is copying the output PDF from the build directory
to the local working directory.&lt;&#x2F;p&gt;
&lt;p&gt;If you want to continue using your standard LaTeX build tool locally (in my case this is &lt;code&gt;latexmk&lt;&#x2F;code&gt;),
you can check whether the &lt;code&gt;TRAVIS&lt;&#x2F;code&gt; environment variable is defined enabling separate build
processes as in the example compilation step below.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;makefile&quot; class=&quot;language-makefile &quot;&gt;&lt;code class=&quot;language-makefile&quot; data-lang=&quot;makefile&quot;&gt;%.pdf: %.tex | $(build_dir)
ifdef TRAVIS
	tectonic -o $(build_dir) --keep-intermediates -r0 $&amp;lt;
	# Run biber if we find a .bcf file in the output
	if [ -f $(build_dir)&amp;#x2F;$(notdir $(&amp;lt;:.tex=.bcf)) ]; then \
		biber --input-directory $(build_dir) $(notdir $(&amp;lt;:.tex=)); \
	fi
	tectonic -o $(build_dir) --keep-intermediates $&amp;lt;
else
	latexmk -outdir=$(build_dir) -pdf -xelatex $&amp;lt;
endif
	cp $(build_dir)&amp;#x2F;$(notdir $@) .
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;The &lt;code&gt;TRAVIS&lt;&#x2F;code&gt; environment variable is defined on all Travis instances and can be set locally for
testing using the command&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;bash&quot; class=&quot;language-bash &quot;&gt;&lt;code class=&quot;language-bash&quot; data-lang=&quot;bash&quot;&gt;TRAVIS=true make
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;This sets the variable &lt;code&gt;TRAVIS&lt;&#x2F;code&gt; just for the single command. Note that you will want to run a
&lt;code&gt;make clean&lt;&#x2F;code&gt; between running with the different build systems as there will be incompatibility
between versions.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;configuring-travis-ci&quot;&gt;Configuring Travis CI&lt;&#x2F;h2&gt;
&lt;p&gt;With a LaTeX distribution that is suitable for CI in &lt;a href=&quot;https:&#x2F;&#x2F;tectonic-typesetting.github.io&quot;&gt;Tectonic&lt;&#x2F;a&gt;, how do I actually use Travis CI?&lt;&#x2F;p&gt;
&lt;p&gt;This is a complicated process, which took me 81 attempts to get right. I hope this guide helps you
in making significantly fewer than me. The outline of the steps is below. The first three are
configuring and linking your accounts on GitHub and Travis. The rest of the steps are explained in
more detail in the rest of this document.&lt;&#x2F;p&gt;
&lt;ol&gt;
&lt;li&gt;Create a public repository on GitHub. Travis CI only works with GitHub and while it does work
with private repositories they requires a paid account with Travis.&lt;&#x2F;li&gt;
&lt;li&gt;Using your GitHub account, sign in to &lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;marketplace&#x2F;travis-ci&#x2F;plan&#x2F;MDIyOk1hcmtldHBsYWNlTGlzdGluZ1BsYW43MA==#pricing-and-setup&quot;&gt;GitHub&lt;&#x2F;a&gt; and add the Travis CI app to the
repository you want to activate. You&#x27;ll need Admin permissions for that repository.&lt;&#x2F;li&gt;
&lt;li&gt;Once signed in to Travis CI, go to your profile page and enable the repository you want to
build.&lt;&#x2F;li&gt;
&lt;li&gt;Create a &lt;code&gt;.travis.yml&lt;&#x2F;code&gt; file in the repository which tells Travis CI what to do. What you need
to put in the file is addressed [below]({{&amp;lt;ref &amp;quot;#travis.yml&amp;quot; &amp;gt;}}).&lt;&#x2F;li&gt;
&lt;li&gt;Commit the &lt;code&gt;.travis.yml&lt;&#x2F;code&gt; file to the repository and push to GitHub. Travis will see the commit
and start the build process.&lt;&#x2F;li&gt;
&lt;li&gt;Any further commits to the repository, whether to the master branch, other branches, tags, or
pull-requests will trigger a build on Travis.&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;h3 id=&quot;travis.yml&quot;&gt;Creating a .travis.ml file&lt;&#x2F;h3&gt;
&lt;p&gt;The &lt;code&gt;.travis.yml&lt;&#x2F;code&gt; file is comprised of a number of sections which I have described individually
below. The complete file is available for &lt;a href=&quot;&#x2F;code&#x2F;travis_latex&#x2F;.travis.yml&quot;&gt;download&lt;&#x2F;a&gt;. To get Travis to start building
your repository, commit the &lt;code&gt;.travis.yml&lt;&#x2F;code&gt; file to the repository and push the commit to GitHub.&lt;&#x2F;p&gt;
&lt;p&gt;The first part of the &lt;code&gt;.travis.yml&lt;&#x2F;code&gt; file is specifying the &lt;a href=&quot;https:&#x2F;&#x2F;docs.travis-ci.com&#x2F;user&#x2F;languages&#x2F;&quot;&gt;language&lt;&#x2F;a&gt;, which is
choosing which of the base containers to use for the compilation.&lt;sup class=&quot;footnote-reference&quot;&gt;&lt;a href=&quot;#3&quot;&gt;2&lt;&#x2F;a&gt;&lt;&#x2F;sup&gt; Since we are using conda for
installing any dependencies we can use the &lt;code&gt;minimal&lt;&#x2F;code&gt; image.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;yaml&quot; class=&quot;language-yaml &quot;&gt;&lt;code class=&quot;language-yaml&quot; data-lang=&quot;yaml&quot;&gt;language: minimal
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;The next part is the cache, where I added the &lt;code&gt;$HOME&#x2F;.cache&#x2F;Tectonic&lt;&#x2F;code&gt; directory. This is where
tectonic stores the downloaded tex packages. Rather than downloading all the packages required on
every build, the state of this entire directory will be downloaded at the start of the build and
updated when changed at the end of the build. This significantly speeds up the build process.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;yaml&quot; class=&quot;language-yaml &quot;&gt;&lt;code class=&quot;language-yaml&quot; data-lang=&quot;yaml&quot;&gt;cache:
  directories:
    - $HOME&amp;#x2F;.cache&amp;#x2F;Tectonic
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;An additional directory that can be cached is &lt;code&gt;$HOME&#x2F;miniconda&lt;&#x2F;code&gt; so the conda packages are also
pre-installed on every build. This is less of a speed-up, although it can be helpful.&lt;&#x2F;p&gt;
&lt;p&gt;The next step is specifying the steps to occur &lt;code&gt;before_install&lt;&#x2F;code&gt;. Travis has a number of steps in the
&lt;a href=&quot;https:&#x2F;&#x2F;docs.travis-ci.com&#x2F;user&#x2F;customizing-the-build&#x2F;#The-Build-Lifecycle&quot;&gt;lifecycle of a build&lt;&#x2F;a&gt; providing ways of breaking the build into logical
steps. Each item in the list below is a bash command, which is run to update the environment of the
test container. This downloads and installs both &lt;code&gt;biber&lt;&#x2F;code&gt; and conda, with conda then being used to
install tectonic. Any other dependencies that are required can also be added to this step, say if
you are converting Markdown to LaTeX with pandoc you could add &lt;code&gt;conda install pandoc&lt;&#x2F;code&gt;.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;yaml&quot; class=&quot;language-yaml &quot;&gt;&lt;code class=&quot;language-yaml&quot; data-lang=&quot;yaml&quot;&gt;before_install:
  # Download and install conda
  - wget https:&amp;#x2F;&amp;#x2F;repo.continuum.io&amp;#x2F;miniconda&amp;#x2F;Miniconda3-latest-Linux-x86_64.sh -O $HOME&amp;#x2F;miniconda.sh
  - bash $HOME&amp;#x2F;miniconda.sh -b -u -p $HOME&amp;#x2F;miniconda
  - export PATH=&amp;quot;$HOME&amp;#x2F;miniconda&amp;#x2F;bin:$PATH&amp;quot;
  - hash -r

    # Install tectonic
  - conda install -y -c conda-forge tectonic==0.1.8
  - conda install -y -c malramsay biber==2.5
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;The final step is the &lt;code&gt;script&lt;&#x2F;code&gt;, the code that is used to determine success or failure. Like the
&lt;code&gt;before_install&lt;&#x2F;code&gt; section, this is a list of commands which are executed one after another. Unlike the
&lt;code&gt;before_install&lt;&#x2F;code&gt; section all these commands are run, even when a command fails.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;yaml&quot; class=&quot;language-yaml &quot;&gt;&lt;code class=&quot;language-yaml&quot; data-lang=&quot;yaml&quot;&gt;script:
  - make
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;The script section is also where you can add any other checks, like ensuring you haven&#x27;t left in any
TODOs, or spelling mistakes. You can run any code you like and if the exit code is zero it is deemed successful,
while a non-zero exit code is a build failure. Note that a failing build will not go on to produce
a release.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;deploying-to-github-releases&quot;&gt;Deploying to GitHub Releases&lt;&#x2F;h2&gt;
&lt;p&gt;Having configured Travis to compile our document on every commit, it would be nice to actually do
something with the resulting document. Every repository on GitHub has releases, which can be
accessed by clicking on the releases link which is outlined in red on the image below.&lt;&#x2F;p&gt;
&lt;p&gt;&lt;img src=&quot;&#x2F;code&#x2F;latex_travis&#x2F;releases.jpg&quot; alt=&quot;The link to the releases page on GitHub&quot; &#x2F;&gt;&lt;&#x2F;p&gt;
&lt;p&gt;GitHub automatically creates a release for every tagged commit in the repository, creating a
downloadable &lt;code&gt;.zip&lt;&#x2F;code&gt; and &lt;code&gt;.tar.gz&lt;&#x2F;code&gt; of the state of the repository at that commit. It is also possible
to edit each of the releases, adding release notes or additional files like installers for a
variety of platforms. In this case we are going to use the GitHub releases to store the compiled
document for each tagged release providing a historical view of the document which is linked to the
code generating it.&lt;&#x2F;p&gt;
&lt;p&gt;In writing of my PhD thesis it makes sense to tag releases using &lt;a href=&quot;https:&#x2F;&#x2F;semver.org&#x2F;&quot;&gt;Semantic Versioning&lt;&#x2F;a&gt;, or at least
&lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;malramsay64&#x2F;phd-thesis&#x2F;blob&#x2F;master&#x2F;planning&#x2F;versioning.md&quot;&gt;a version of it&lt;&#x2F;a&gt;. Other documents are much less linear, it might make sense to tag a
talk with the location it will be given, or the name of the conference. The requirements are
basically use numbers, letters and any of &lt;code&gt;._-+&#x2F;&lt;&#x2F;code&gt; --- see the &lt;a href=&quot;https:&#x2F;&#x2F;git-scm.com&#x2F;docs&#x2F;git-check-ref-format&quot;&gt;git-check-ref-format&lt;&#x2F;a&gt; documentation for more
specific details).&lt;sup class=&quot;footnote-reference&quot;&gt;&lt;a href=&quot;#2&quot;&gt;3&lt;&#x2F;a&gt;&lt;&#x2F;sup&gt;&lt;&#x2F;p&gt;
&lt;p&gt;You can create a tag &lt;code&gt;my_tag&lt;&#x2F;code&gt; for a release by running the command&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;bash&quot; class=&quot;language-bash &quot;&gt;&lt;code class=&quot;language-bash&quot; data-lang=&quot;bash&quot;&gt;git tag -a my_tag
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;which will open an editor to write a message. A typical tag message is the repository name followed
by the tag, although that isn&#x27;t the only approach. Like commit messages, you can specify
the message using the &lt;code&gt;-m&lt;&#x2F;code&gt; option.&lt;&#x2F;p&gt;
&lt;p&gt;By default, git doesn&#x27;t push tags to a remote, requiring the
&lt;code&gt;--tags&lt;&#x2F;code&gt; option&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;bash&quot; class=&quot;language-bash &quot;&gt;&lt;code class=&quot;language-bash&quot; data-lang=&quot;bash&quot;&gt;git push --tags
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Before pushing your newly created tag, you are going to want to configure Travis to upload releases
to GitHub. The best method for this is to use the Travis command line client, which can be installed
by following &lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;travis-ci&#x2F;travis.rb#installation&quot;&gt;these instructions&lt;&#x2F;a&gt;. Once installed you can run the command&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;bash&quot; class=&quot;language-bash &quot;&gt;&lt;code class=&quot;language-bash&quot; data-lang=&quot;bash&quot;&gt;travis setup releases
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;which will prompt for your GitHub credentials and other information about the release. These are
used to generate a personal access token for GitHub which Travis uses to authenticate when uploading
the release. The &lt;code&gt;travis&lt;&#x2F;code&gt; client will encrypt the token, and update your &lt;code&gt;.travis.yml&lt;&#x2F;code&gt; with a deploy
section which looks something like this;&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;yaml&quot; class=&quot;language-yaml &quot;&gt;&lt;code class=&quot;language-yaml&quot; data-lang=&quot;yaml&quot;&gt;deploy:
  provider: releases
  api_key:
    secure: # your encrypted token will be here
  file: thesis.pdf
  skip_cleanup: true
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Since we only want to deploy on tagged commits, we can use the
&lt;a href=&quot;https:&#x2F;&#x2F;docs.travis-ci.com&#x2F;user&#x2F;deployment#Conditional-Releases-with-on%3A&quot;&gt;conditional deployment&lt;&#x2F;a&gt; options to
conditionally deploy. This gives the following deploy section.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;yaml&quot; class=&quot;language-yaml &quot;&gt;&lt;code class=&quot;language-yaml&quot; data-lang=&quot;yaml&quot;&gt;deploy:
  provider: releases
  api_key:
    secure: # your encrypted token will be here
  file: thesis.pdf
  skip_cleanup: true
  on:
    tags: true
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;h2 id=&quot;conclusion&quot;&gt;Conclusion&lt;&#x2F;h2&gt;
&lt;p&gt;I have made the entire &lt;code&gt;.travis.yml&lt;&#x2F;code&gt; file available for &lt;a href=&quot;&#x2F;code&#x2F;travis_latex&#x2F;.travis.yml&quot;&gt;download&lt;&#x2F;a&gt; should you want to
get started quickly. Or alternatively have a look at my repositories &lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;malramsay64&#x2F;usyd-beamer-theme&quot;&gt;usyd-beamer-theme&lt;&#x2F;a&gt; or
&lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;malramsay64&#x2F;phd-thesis&quot;&gt;phd-thesis&lt;&#x2F;a&gt; which I have set up to use this workflow. This file along with the &lt;a href=&quot;&#x2F;code&#x2F;travis_latex&#x2F;Makefile&quot;&gt;Makefile&lt;&#x2F;a&gt; this
should enable this process to work for the compilation of most documents.&lt;&#x2F;p&gt;
&lt;p&gt;While the process is complicated to set up, once completed it shouldn&#x27;t require much effort to
maintain. There is the rationale it might save you some time, however I think it is cool which is
all justification I needed.&lt;&#x2F;p&gt;
&lt;hr &#x2F;&gt;
&lt;p&gt;&lt;strong&gt;Update 2018-07-26:&lt;&#x2F;strong&gt; Since initially writing this post I have created a conda package to
distribute the biber binary. I have updated the post and included files to use this package since it
significantly simplifies the installation. The steps from the original &lt;code&gt;travis.yml&lt;&#x2F;code&gt; file to manually
install biber are below.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;yaml&quot; class=&quot;language-yaml &quot;&gt;&lt;code class=&quot;language-yaml&quot; data-lang=&quot;yaml&quot;&gt;  # Download and install biber installing executable as biber2.5
  - wget https:&amp;#x2F;&amp;#x2F;sourceforge.net&amp;#x2F;projects&amp;#x2F;biblatex-biber&amp;#x2F;files&amp;#x2F;biblatex-biber&amp;#x2F;2.5&amp;#x2F;binaries&amp;#x2F;Linux&amp;#x2F;biber-linux_x86_64.tar.gz -O $HOME&amp;#x2F;downloads&amp;#x2F;biber.tar.gz
  - tar xvzf $HOME&amp;#x2F;downloads&amp;#x2F;biber.tar.gz -C $HOME&amp;#x2F;bin
  - mv $HOME&amp;#x2F;bin&amp;#x2F;biber $HOME&amp;#x2F;bin&amp;#x2F;biber2.5
  - export PATH=&amp;quot;$HOME&amp;#x2F;bin:$PATH&amp;quot;
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;div class=&quot;footnote-definition&quot; id=&quot;1&quot;&gt;&lt;sup class=&quot;footnote-definition-label&quot;&gt;1&lt;&#x2F;sup&gt;
&lt;p&gt;I should note that MiKTeX will install &lt;code&gt;biber&lt;&#x2F;code&gt; on a Windows system. So if you wanted to set
up a Windows CI config I guess MiKTeX is a great approach.&lt;&#x2F;p&gt;
&lt;&#x2F;div&gt;
&lt;div class=&quot;footnote-definition&quot; id=&quot;2&quot;&gt;&lt;sup class=&quot;footnote-definition-label&quot;&gt;3&lt;&#x2F;sup&gt;
&lt;p&gt;There are more special characters supported, I just listed the most common ones. See the
&lt;a href=&quot;https:&#x2F;&#x2F;git-scm.com&#x2F;docs&#x2F;git-check-ref-format&quot;&gt;docs&lt;&#x2F;a&gt; for more specific information.&lt;&#x2F;p&gt;
&lt;&#x2F;div&gt;
&lt;div class=&quot;footnote-definition&quot; id=&quot;3&quot;&gt;&lt;sup class=&quot;footnote-definition-label&quot;&gt;2&lt;&#x2F;sup&gt;
&lt;p&gt;You may notice that the minimal is not listed in the documentation. There is a
&lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;travis-ci&#x2F;docs-travis-ci-com&#x2F;issues&#x2F;910#issuecomment-356915625&quot;&gt;GitHub issues&lt;&#x2F;a&gt; to
rectify this in which there is documentation.&lt;&#x2F;p&gt;
&lt;&#x2F;div&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>Experi: A tool for computational experiments</title>
        <published>2018-06-25T00:00:00+00:00</published>
        <updated>2018-06-25T00:00:00+00:00</updated>
        <author>
          <name>Unknown</name>
        </author>
        <link rel="alternate" href="https://malramsay.com/post/experi-a-tool-for-computational-experiments/" type="text/html"/>
        <id>https://malramsay.com/post/experi-a-tool-for-computational-experiments/</id>
        
        <content type="html">&lt;h2 id=&quot;abstract&quot;&gt;Abstract&lt;&#x2F;h2&gt;
&lt;p&gt;One of the key features of computational experiments is being able to run the experiment over
a large variable space. However, in my experience there aren&#x27;t tools available to assist with this,
particularly in the realm of High Performance Computing (HPC), where bash arrays and loops are
commonplace. Using the current toolset, I made lots of errors in the specification of files,
turning a &#x27;quick edit&#x27; into a tedious process of find the bug. To make complicated experiment
variable expression a simple and intuitive task I have developed &lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;malramsay64&#x2F;experi&quot;&gt;Experi&lt;&#x2F;a&gt;, which takes a list of
commands and variables from a &lt;a href=&quot;https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;YAML&quot;&gt;YAML&lt;&#x2F;a&gt; input file and either runs them for testing or creates and
submits files to an HPC scheduler. I have designed Experi to be a simple replacement for current
workflows, using shell commands into which you add variables, and variables defined along with how
they are combined. I have been using Experi as I have been developing it and though it hasn&#x27;t
prevented me from running software with bugs, I can specify the variables I want to use for
a simulation without having to worry about the execution. As an added benefit, I also have
a complete record of every simulation that I run which is version controllable.&lt;&#x2F;p&gt;
&lt;p&gt;Experi is installable using both &lt;code&gt;pip&lt;&#x2F;code&gt; and &lt;code&gt;conda&lt;&#x2F;code&gt;,&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$ pip install experi
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Currently, Experi only runs on python 3.6, because that is what I use and &amp;quot;premature optimisation is
the root of all evil.&amp;quot; If you would like to use Experi and can&#x27;t use python 3.6, please let me know
along with which python version you do use and I can look into supporting it.&lt;&#x2F;p&gt;
&lt;hr &#x2F;&gt;
&lt;p&gt;I work in science where most compute intensive problems are handled by High Performance Computing
(HPC) clusters. From the user perspective the entire cluster appears as a single host, with
resources accessible through a job scheduler. You submit work as a script to the scheduler along
with the resources required, where it will wait in a queue for those resources to be available. Once
your script is running, it can access the resources allocated. When using a parallelism technology
like MPI, the MPI software will detect all the resources and connecting the appropriate pieces
together. This makes for a reasonably simple method of accessing the vast compute capabilities of
HPCs.&lt;&#x2F;p&gt;
&lt;p&gt;Bash scripts are normally used to specify the work to take place; with additional scheduler
directives denoting the resources required. One of scheduler directives you can specify is to make
the job as an array job, where multiple instances of your file are submitted at once with an array
index. In some ways this is similar to a numpy array operation rather than a &lt;code&gt;for&lt;&#x2F;code&gt; loop. The array
job is able to make best use of the scheduler, splitting a large task into many smaller tasks
which are easier to fit onto the cluster. This kind of workflow is ideal for running an experiment
at different conditions, or with many replications.&lt;&#x2F;p&gt;
&lt;p&gt;With the scheduler I use, &lt;a href=&quot;http:&#x2F;&#x2F;www.pbspro.org&#x2F;&quot;&gt;PBSPro&lt;&#x2F;a&gt; an array job is specified using the &lt;code&gt;-J&lt;&#x2F;code&gt; flag and takes an
argument instructing how to create the values. The argument has the form
&lt;code&gt;&amp;lt;start_index&amp;gt;-&amp;lt;end_index&amp;gt;:&amp;lt;increment&amp;gt;&lt;&#x2F;code&gt;, e.g. &lt;code&gt;1-10:2&lt;&#x2F;code&gt; will generate five jobs for the scheduler
which will each have the bash variable &lt;code&gt;PBS_ARRAY_INDEX&lt;&#x2F;code&gt; containing one of &lt;code&gt;1,3,5,7,9&lt;&#x2F;code&gt;. While this
can be useful where you need to vary a parameter over a sequence of integers, the more likely use
case is creating a list of values which you are then indexing using &lt;code&gt;PBS_ARRAY_INDEX&lt;&#x2F;code&gt; like the
example below&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;bash&quot; class=&quot;language-bash &quot;&gt;&lt;code class=&quot;language-bash&quot; data-lang=&quot;bash&quot;&gt;ARRAY = (1.0 1.5 2.0 2.5 3.0 4.0)
value = &amp;quot;${ARRAY[$PBS_ARRAY_INDEX]}&amp;quot;
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;When specifying my jobs for the scheduler I was making lots of mistakes writing a job script that
was valid and reflected my intentions. This was difficult in that each job involved changing
multiple variables and I only had a single index. While Bash is excellent for specifying complex
commands for the computer to interpret, when dealing with long commands and lots of variables it
quickly becomes too complicated to easily read. When I am designing a computational experiment
I think about the commands I want to run separately from the variables I am going to use in those
commands.&lt;&#x2F;p&gt;
&lt;p&gt;My experiments have a recipe that look something like;&lt;&#x2F;p&gt;
&lt;ol&gt;
&lt;li&gt;Create a state with the properties I desire. This could be a liquid, a crystal, or a combination
of the two.&lt;&#x2F;li&gt;
&lt;li&gt;Bring the initial state to the conditions I want to collect data, ensuring the new state is
representative of those conditions.&lt;&#x2F;li&gt;
&lt;li&gt;Collect lots of data on how my experiment behaves at this specific condition.&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;p&gt;I have presented each of these steps abstractly as they are applicable to many different
experiments. It is also the level of abstraction for the command line interface, &lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;malramsay64&#x2F;statdyn-simulation&quot;&gt;sdrun&lt;&#x2F;a&gt;, I have
written to help run my experiments. A command I would run for the create step would be&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;bash&quot; class=&quot;language-bash &quot;&gt;&lt;code class=&quot;language-bash&quot; data-lang=&quot;bash&quot;&gt;$ sdrun --temperature 0.1 --num-steps 100 --space-group p2 create configuration.out
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;I specify the variables for the simulation, the type of simulation being &lt;code&gt;create&lt;&#x2F;code&gt;, and then the
output file.&lt;&#x2F;p&gt;
&lt;p&gt;For each of the three steps above I now want to define the specific variables I am using;&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;
&lt;p&gt;Crystal Structure:&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;p2&lt;&#x2F;li&gt;
&lt;li&gt;p2gg&lt;&#x2F;li&gt;
&lt;li&gt;pg&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;Temperature:&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;0.4&lt;&#x2F;li&gt;
&lt;li&gt;0.45&lt;&#x2F;li&gt;
&lt;li&gt;0.50&lt;&#x2F;li&gt;
&lt;li&gt;0.6&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;&#x2F;li&gt;
&lt;li&gt;
&lt;p&gt;Steps:&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;1_000_000&lt;&#x2F;li&gt;
&lt;li&gt;1_000_000&lt;&#x2F;li&gt;
&lt;li&gt;500_000&lt;&#x2F;li&gt;
&lt;li&gt;100_000&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;I want to run the same simulation for each of three different crystal
structures, p2, p2gg, and pg. This simulation will be over a range of
temperatures, with smaller spacing at lower temperatures where there is
interesting behaviour. Also the effects I am looking for happen faster at
higher temperatures so I am going to run shorter simulations for the higher
temperatures.&lt;&#x2F;p&gt;
&lt;p&gt;In presenting the values for the variables in a concise way, a table is possibly clearer though much
less concise, I am nearly using &lt;a href=&quot;https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;YAML&quot;&gt;YAML&lt;&#x2F;a&gt; syntax. YAML is a method of encoding both structure and
data in a human readable format. You can think about it as a way of defining python lists and dicts
in a text file. Lists are a series of bullet points,&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;yaml&quot; class=&quot;language-yaml &quot;&gt;&lt;code class=&quot;language-yaml&quot; data-lang=&quot;yaml&quot;&gt;- 0
- 1
- 2
- 3
- 4
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;with the snippet above being the list &lt;code&gt;[0, 1, 2, 3, 4]&lt;&#x2F;code&gt; in python. Dicts are
mappings of values using the colon &lt;code&gt;:&lt;&#x2F;code&gt; as a separator&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;yaml&quot; class=&quot;language-yaml &quot;&gt;&lt;code class=&quot;language-yaml&quot; data-lang=&quot;yaml&quot;&gt;hello: world
foo: bar
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;being the dict &lt;code&gt;{&amp;quot;hello&amp;quot;: &amp;quot;world&amp;quot;, &amp;quot;foo&amp;quot;: &amp;quot;bar&amp;quot;}&lt;&#x2F;code&gt; in python.&lt;&#x2F;p&gt;
&lt;p&gt;With the goal of being able to easily specify commands required for an experiment in addition to the
variables to fill into those commands, I have developed &lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;malramsay64&#x2F;experi&quot;&gt;Experi&lt;&#x2F;a&gt;. It takes a YAML file as input,
declaring the experiment that will take place, and either creates and submits files to the scheduler
or runs the experiment in the current shell for testing. Taking the example I constructed above, to
completely specify the experiment for Experi;&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;yaml&quot; class=&quot;language-yaml &quot;&gt;&lt;code class=&quot;language-yaml&quot; data-lang=&quot;yaml&quot;&gt;# experiment.yml
commands:
  - sdrun --space-group {crystal} -t 0.3 -s 1000 create configuration-{crystal}-0.3.out
  - &amp;gt;
    sdrun
    --space-group {crystal}
    --init-temp 0.3
    --temperature {temperature}
    --num-steps {steps}
    equil
    config-{crystal}-0.3.out
    config-{crystal}-{temperature:.2f}.out
  - sdrun --space-group {crystal} -t {temperature} -s {steps} prod config-{crystal}-{temperature:.2f}.out

variables:
  crystal:
    - p2
    - pg
    - p2gg

  zip:
    temperature:
      - 0.4
      - 0.45
      - 0.50
      - 0.6
    steps:
      - 1_000_000
      - 1_000_000
      - 500_000
      - 100_000

pbs:
  ncpus: 12
  walltime: 100:00:00
  j: oe
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;There are three commands which will be run; firstly a &lt;code&gt;create&lt;&#x2F;code&gt; step, secondly an &lt;code&gt;equil&lt;&#x2F;code&gt;
step, finally a &lt;code&gt;prod&lt;&#x2F;code&gt; step. The longer command of the &lt;code&gt;equil&lt;&#x2F;code&gt; step is broken
over multiple lines, and since I am using the YAML syntax for a multiline string &lt;code&gt;&amp;gt;&lt;&#x2F;code&gt;, it is
interpreted as a single long string hence no need for line continuation characters. Where I want to
replace the variables I have enclosed the variable name in braces (&lt;code&gt;{}&lt;&#x2F;code&gt;), which will be
automatically replaced using &lt;a href=&quot;https:&#x2F;&#x2F;pyformat.info&#x2F;&quot;&gt;python string formatting&lt;&#x2F;a&gt;. Using this format allows for modifiers
like &lt;code&gt;{temperature:.2f}&lt;&#x2F;code&gt; which is replaced with the temperature as a float having two decimal
places, i.e. &lt;code&gt;0.4&lt;&#x2F;code&gt; -&amp;gt; &lt;code&gt;&#x27;0.40&#x27;&lt;&#x2F;code&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;Commands in Experi are guaranteed to be run in a linear fashion, i.e. none of the &lt;code&gt;equil&lt;&#x2F;code&gt;
simulations will run until the &lt;code&gt;create&lt;&#x2F;code&gt; simulations are finished, with the additional constraint
that all previous steps are successful. This holds for both running in the local terminal, or using
the job scheduler.&lt;&#x2F;p&gt;
&lt;p&gt;Variables for the experiment are specified in their own section, with the variable name as the
dictionary key, followed by a list of values---a single value is also supported. Variables can take
on any valid string with the exception of &lt;code&gt;zip&lt;&#x2F;code&gt; or &lt;code&gt;product&lt;&#x2F;code&gt; which are used define how the variables
are combined. &lt;code&gt;zip&lt;&#x2F;code&gt; acts on the variables like a zipper, taking variables with the same number of
values and matching the values up 1:1. In this case we get&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;temperature: 0.4 steps: 1_000_000&lt;&#x2F;li&gt;
&lt;li&gt;temperature: 0.45 steps: 1_000_000&lt;&#x2F;li&gt;
&lt;li&gt;temperature: 0.50 steps: 500_000&lt;&#x2F;li&gt;
&lt;li&gt;temperature: 0.6 steps: 100_000&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;&lt;code&gt;product&lt;&#x2F;code&gt; is also somewhat self-descriptive, taking the &#x27;product&#x27; of all the values to give all
possible combinations. &lt;code&gt;product&lt;&#x2F;code&gt; is the default operation and is not explicitly required so each of
the three crystals will have simulations with the temperature and steps listed above.&lt;&#x2F;p&gt;
&lt;p&gt;Where multiple &lt;code&gt;zip&lt;&#x2F;code&gt; operations are required, a list of dictionaries can be supplied to the &lt;code&gt;zip&lt;&#x2F;code&gt;
key with each being zipped separately before taking the product of each item in the list.&lt;&#x2F;p&gt;
&lt;p&gt;The final part of the YAML file is the specification of the options for the PBS scheduler. There are
some values which can be passed as keys;&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;select&#x2F;nodes&lt;&#x2F;li&gt;
&lt;li&gt;ncpus&#x2F;cpus&lt;&#x2F;li&gt;
&lt;li&gt;ngpus&#x2F;gpus&lt;&#x2F;li&gt;
&lt;li&gt;memory&#x2F;mem&lt;&#x2F;li&gt;
&lt;li&gt;walltime&lt;&#x2F;li&gt;
&lt;li&gt;cputime&lt;&#x2F;li&gt;
&lt;li&gt;project&lt;&#x2F;li&gt;
&lt;li&gt;setup&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;While any options not specifically encoded can be passed by using the flag as the key. In the
example above I have used &lt;code&gt;j: oe&lt;&#x2F;code&gt; which will become the scheduler option &lt;code&gt;-j oe&lt;&#x2F;code&gt; which is specifying
that the stdout (&lt;code&gt;o&lt;&#x2F;code&gt;) and stderr (&lt;code&gt;e&lt;&#x2F;code&gt;) will be joined in the stdout stream.&lt;&#x2F;p&gt;
&lt;p&gt;The &lt;code&gt;setup&lt;&#x2F;code&gt; key allows for commands to load the required variables, modules, or anything else that
is required before each command is run.&lt;&#x2F;p&gt;
&lt;p&gt;It should be noted that YAML is not without &lt;a href=&quot;https:&#x2F;&#x2F;arp242.net&#x2F;weblog&#x2F;yaml_probably_not_so_great_after_all.html&quot;&gt;idiosyncrasies&lt;&#x2F;a&gt;. However, it is
powerful and widespread and I like and understand the format. The structure of Experi is not
specific to YAML and could easily work with JSON or TOML input files if someone wants to put effort
into the implementation and documentation.&lt;&#x2F;p&gt;
&lt;p&gt;With the input file specified, assuming it is named &lt;code&gt;experiment.yml&lt;&#x2F;code&gt;, the simulation can be
submitted to the scheduler using the command&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;bash&quot; class=&quot;language-bash &quot;&gt;&lt;code class=&quot;language-bash&quot; data-lang=&quot;bash&quot;&gt;$ experi
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;which will submit the job to the scheduler if it is available. Regardless of the presence of the
scheduler, a file is created for each command which includes all the combinations of variables
specified in a bash array which is suitable for manual submission.&lt;&#x2F;p&gt;
&lt;p&gt;Using Experi for the design and execution of computational experiments has significantly benefited
my science workflow. This is most noticeable with the replication and modification of experiments.
I have a complete record of all the commands and options required for the experiment to run and
I can focus on modifying values without having to worry about how to loop over the range of
variables I require.&lt;&#x2F;p&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>Distributing a Hoomd Plugin</title>
        <published>2018-06-03T00:00:00+00:00</published>
        <updated>2018-06-03T00:00:00+00:00</updated>
        <author>
          <name>Unknown</name>
        </author>
        <link rel="alternate" href="https://malramsay.com/post/distributing-a-hoomd-plugin/" type="text/html"/>
        <id>https://malramsay.com/post/distributing-a-hoomd-plugin/</id>
        
        <content type="html">&lt;p&gt;A piece of software I have been using in my research is &lt;a href=&quot;http:&#x2F;&#x2F;hoomd-blue.readthedocs.io&#x2F;en&#x2F;stable&#x2F;&quot;&gt;Hoomd&lt;&#x2F;a&gt;,
a &#x27;relatively&#x27; new package for running Molecular Dynamics (MD) simulations.
These MD simulations have the basic premise of
throwing hundreds of balls into a box and shaking it
to find out what happens.
The relative newness of Hoomd
is in comparison to other software packages like &lt;a href=&quot;http:&#x2F;&#x2F;lammps.sandia.gov&#x2F;&quot;&gt;LAMMPS&lt;&#x2F;a&gt; and &lt;a href=&quot;http:&#x2F;&#x2F;www.gromacs.org&#x2F;&quot;&gt;GROMACS&lt;&#x2F;a&gt;
which have been around for decades,
while the initial release of Hoomd was in 2012.
There are some major benefits of a newer approach to MD simulations
and Hoomd is most notably designed to leverage the computational power of GPUs.
Despite the benefits of a modern approach,
Hoomd doesn&#x27;t have the range of built-in simulation types
of the more mature software packages.
To allow researchers with specific problems to still use Hoomd,
it has a plugin architecture
to simplify the implementation of custom functionality.&lt;&#x2F;p&gt;
&lt;p&gt;This article documents how I took a plugin I wrote for my research
and made it installable in a conda environment alongside Hoomd.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;the-plugin&quot;&gt;The Plugin&lt;&#x2F;h2&gt;
&lt;p&gt;The example plugin is one which I have written for my own research.
It implements a harmonic pinning potential on the
positions and rotations of rigid bodies
in a MD simulation.
The complete plugin is available for reference on &lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;malramsay64&#x2F;hoomd-harmonic-force&quot;&gt;GitHub&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;To get started writing your own plugin
there is a short guide in the &lt;a href=&quot;http:&#x2F;&#x2F;hoomd-blue.readthedocs.io&#x2F;en&#x2F;stable&#x2F;developer.html&quot;&gt;documentation&lt;&#x2F;a&gt;
which directs you to the &lt;code&gt;example_plugin&lt;&#x2F;code&gt; in the Hoomd source code.
This code is located on &lt;a href=&quot;https:&#x2F;&#x2F;bitbucket.org&#x2F;glotzer&#x2F;hoomd-blue&#x2F;src&#x2F;maint&#x2F;&quot;&gt;bitbucket&lt;&#x2F;a&gt; and there is also a copy on &lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;joaander&#x2F;hoomd-blue&quot;&gt;GitHub&lt;&#x2F;a&gt;.
I also found that finding a class that shared the same base class was useful for writing the implementation.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;the-makefile&quot;&gt;The Makefile&lt;&#x2F;h2&gt;
&lt;p&gt;The example code for the &lt;code&gt;example_plugin&lt;&#x2F;code&gt; doesn&#x27;t include a &lt;code&gt;Makefile&lt;&#x2F;code&gt;,
requiring multiple steps to compile the plugin.
A Makefile is useful in this case for
both being able to install with a single command
and ensuring the install options are consistent with the conda installed Hoomd.&lt;&#x2F;p&gt;
&lt;p&gt;While my plugin doesn&#x27;t use CUDA for calculations,
for compatibility with the conda installed Hoomd
enabling CUDA on Linux is required.
When I was compiling with CUDA disabled I came across unusual behaviour,
for example numbers being far too large, which appeared to be a buffer overflow.
I believe this is because of conditional definitions in the header files
resulting in a slightly different memory layout
for a class defined with or without CUDA.&lt;&#x2F;p&gt;
&lt;p&gt;If you are able to build your plugin using &lt;code&gt;cmake&lt;&#x2F;code&gt; like the example plugin,
the following Makefile will make compilation simpler
and set the appropriate flags assuming Hoomd was installed using conda.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;makefile&quot; class=&quot;language-makefile &quot;&gt;&lt;code class=&quot;language-makefile&quot; data-lang=&quot;makefile&quot;&gt;build_dir = build

# Check OS to determine if CUDA is enabled, defaults to True
CUDA_ENABLED=True
UNAME_S := $(shell uname -s)
ifeq ($(UNAME_S),Linux)
	CUDA_ENABLED=True
endif
ifeq ($(UNAME_S),Darwin)
	CUDA_ENABLED=False
endif

all: $(build_dir)
	cd $(build_dir); cmake .. -DENABLE_CUDA=$(CUDA_ENABLED)
	$(MAKE) -C $(build_dir)

install: all
	$(MAKE) -C $(build_dir) install

clean:
	rm -rf $(build_dir)

test:
	pytest

$(build_dir):
	mkdir -p $@

.PHONY: test clean
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Apart from making it simpler to install the package manually,
it also makes it simplifies the build using conda.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;packaging&quot;&gt;Packaging&lt;&#x2F;h2&gt;
&lt;p&gt;The conda package manager is the recommended method of installing Hoomd,
primarily because of the simplicity of installation.
Conda allows you to upload your own packages for anyone to download,
which is how we are going to make this plugin simple to install.&lt;&#x2F;p&gt;
&lt;p&gt;To define the package for upload to the &lt;a href=&quot;https:&#x2F;&#x2F;anaconda.org&#x2F;&quot;&gt;Anaconda Cloud&lt;&#x2F;a&gt; repository
we have to write a &lt;code&gt;meta.yaml&lt;&#x2F;code&gt; file.
This defines everything required to build the package including;
the package details,
where to find the source code,
dependencies required to build the package,
dependencies to install the package, and
how to build the package.
There is [extensive documentation][meta.yaml documentation] for the &lt;code&gt;meta.yaml&lt;&#x2F;code&gt; file
which covers a variety of use cases.
The file uses the &lt;a href=&quot;http:&#x2F;&#x2F;yaml.org&#x2F;&quot;&gt;yaml&lt;&#x2F;a&gt; syntax,
which is a method of encoding python data structures
in a more human readable format.&lt;&#x2F;p&gt;
&lt;p&gt;The &lt;code&gt;meta.yaml&lt;&#x2F;code&gt; file contains a series of sections
each containing different information about the package.
Some of these keys are;&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;package&lt;&#x2F;code&gt;,&lt;&#x2F;li&gt;
&lt;li&gt;&lt;code&gt;about&lt;&#x2F;code&gt;,&lt;&#x2F;li&gt;
&lt;li&gt;&lt;code&gt;source&lt;&#x2F;code&gt;,&lt;&#x2F;li&gt;
&lt;li&gt;&lt;code&gt;dependencies&lt;&#x2F;code&gt;, and&lt;&#x2F;li&gt;
&lt;li&gt;&lt;code&gt;test&lt;&#x2F;code&gt;.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;The documentation link above contains more extensive information on the options available for each
of these sections. I am going to explain how I have used all the different sections to create this
plugin.&lt;&#x2F;p&gt;
&lt;p&gt;The &lt;code&gt;package&lt;&#x2F;code&gt; key contains the name and version of the package. This section is
compulsory, while all others are optional, although that doesn&#x27;t mean they
aren&#x27;t required.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;yaml&quot; class=&quot;language-yaml &quot;&gt;&lt;code class=&quot;language-yaml&quot; data-lang=&quot;yaml&quot;&gt;package:
    name: hoomd-harmonic-force
    version: 0.1.7
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;The &lt;code&gt;about&lt;&#x2F;code&gt; key is where you can find more information on the package, including the
homepage and the license the package uses. In this case the homepage is a link to the Github
repository.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;yaml&quot; class=&quot;language-yaml &quot;&gt;&lt;code class=&quot;language-yaml&quot; data-lang=&quot;yaml&quot;&gt;about:
  home: https:&amp;#x2F;&amp;#x2F;github.com&amp;#x2F;malramsay64&amp;#x2F;hoomd-harmonic-force
  license: MIT
  license_file: LICENSE
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;The &lt;code&gt;source&lt;&#x2F;code&gt; section defines where to find the source code for the build phase. I have specified
a specific tag from the git repository, enabling me to checkout an old commit to build an older
version.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;yaml&quot; class=&quot;language-yaml &quot;&gt;&lt;code class=&quot;language-yaml&quot; data-lang=&quot;yaml&quot;&gt;source:
  git_url: https:&amp;#x2F;&amp;#x2F;github.com&amp;#x2F;malramsay64&amp;#x2F;hoomd-harmonic-force.git
  git_rev: v0.1.7
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Another useful option for the &lt;code&gt;source&lt;&#x2F;code&gt; section is use the &lt;code&gt;path&lt;&#x2F;code&gt; key. This key
allows for the specification of the relative (or absolute) path to the source
code. This is most useful in the development process, allowing for quickly
testing whether a change has worked.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;yaml&quot; class=&quot;language-yaml &quot;&gt;&lt;code class=&quot;language-yaml&quot; data-lang=&quot;yaml&quot;&gt;source:
    path: .&amp;#x2F;
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;The next section is the &lt;code&gt;requirements&lt;&#x2F;code&gt;, defining both the packages required for
the build phase, and those required when the package is installed. These package definitions
include the specification of compatible versions of the dependencies. The version specification
is flexible enough for complex build processes, while also having the option for simplicity.
In this guide I describe the simplest method of version specification. When you need something more complex,
the documentation on &lt;a href=&quot;https:&#x2F;&#x2F;conda.io&#x2F;docs&#x2F;user-guide&#x2F;tasks&#x2F;build-packages&#x2F;variants.html&quot;&gt;build variants&lt;&#x2F;a&gt; covers a wide range of use cases.&lt;&#x2F;p&gt;
&lt;p&gt;The build requirements are all the programs required for the build; in this
case compilation with &lt;code&gt;cmake&lt;&#x2F;code&gt;. Since we are building a python module we require
both python and setuptools. I am using python 3.6, so have set the python
version to &lt;code&gt;3.6.*&lt;&#x2F;code&gt; which means any point release of python 3.6, at the time of
writing being 3.6.5. The numpy and Hoomd dependencies are also requirements of
the build process, for which I have specified the latest versions. For python,
numpy and Hoomd you can specify the versions you use for your work. Where you
use multiple versions, like python 2.7, 3.5, and 3.6, the start of the &lt;a href=&quot;https:&#x2F;&#x2F;conda.io&#x2F;docs&#x2F;user-guide&#x2F;tasks&#x2F;build-packages&#x2F;variants.html&quot;&gt;build
variants documentation&lt;&#x2F;a&gt; has examples to adapt. The next
dependency, cudatoolkit, is not strictly necessary in this case, however before
Hoomd v2.2.5 there was a typo in the Hoomd version specification allowing
a newer incompatible version to be installed. The final requirement, &lt;code&gt;cmake&lt;&#x2F;code&gt;,
builds the package and while you might have it installed on your system
already, someone else may not. It should be noted that the &lt;code&gt;nvcc&lt;&#x2F;code&gt; binary is
also required on linux which I can&#x27;t find a conda package for.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;yaml&quot; class=&quot;language-yaml &quot;&gt;&lt;code class=&quot;language-yaml&quot; data-lang=&quot;yaml&quot;&gt;requirements:
  build:
    - python 3.6.*
    - setuptools
    - numpy 1.14.*
    - hoomd 2.3.*
    - cudatoolkit 8.*
    - cmake &amp;gt;=2.8.0
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;With the build dependencies specified, we need to specify the run dependencies.
Version numbers under the [semantic versioning][] scheme have a format of
&lt;code&gt;&amp;lt;major&amp;gt;.&amp;lt;minor&amp;gt;.&amp;lt;patch&amp;gt;&lt;&#x2F;code&gt;. Version numbers with the same major and minor
numbers are typically compatible which is why I have specified the run
requirements with the same minor version as the build requirements. I have come
across previous versions of Hoomd which are only compatible with a singe patch
version of python. Unfortunately, you can&#x27;t pin the python version to a single patch
as conda build &lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;conda&#x2F;conda-build&#x2F;issues&#x2F;2571&quot;&gt;overrides the pinning&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;yaml&quot; class=&quot;language-yaml &quot;&gt;&lt;code class=&quot;language-yaml&quot; data-lang=&quot;yaml&quot;&gt;requirements:
  run:
    - python 3.6.*
    - numpy 1.14.*
    - hoomd 2.3.*
    - cudatoolkit 8.*
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Having waded through the complications of dependency specification we now need
to tell conda how to build our plugin. This is where creating the Makefile is
useful, we can now use the same commands for build process as for a manual
installation; &lt;code&gt;make clean &amp;amp;&amp;amp; make install&lt;&#x2F;code&gt;. The &lt;code&gt;make clean&lt;&#x2F;code&gt; command removes
the build files when using the &lt;code&gt;path&lt;&#x2F;code&gt; option for the &lt;code&gt;source&lt;&#x2F;code&gt;. The other
element of the build section is the number, incremented when uploading another
package with the same version and reset to 0 on a new version. The number acts
as a sub-version number allowing for small fixes, like correcting the version
pinning, without incrementing the package version number.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;yaml&quot; class=&quot;language-yaml &quot;&gt;&lt;code class=&quot;language-yaml&quot; data-lang=&quot;yaml&quot;&gt;build:
  script: make clean &amp;amp;&amp;amp; make install
  number: 0
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;The final section I will discuss is the test section. At its simplest this can
import your newly created package, or it can run a complex unit and integration
test suite. The test section failing will be a fail the build, providing
a final check for bugs before release. In the example below I am checking the
package will import as a first simple test, before running more extensive
testing with &lt;a href=&quot;https:&#x2F;&#x2F;docs.pytest.org&#x2F;en&#x2F;latest&#x2F;&quot;&gt;pytest&lt;&#x2F;a&gt;. To ensure pytest is installed, there is
a &lt;code&gt;requirements&lt;&#x2F;code&gt; key in the &lt;code&gt;test&lt;&#x2F;code&gt; section allowing for the specification
of test specific requirements. I have also specified the files for running the
tests using the &lt;code&gt;source_files&lt;&#x2F;code&gt; key, which is the entire directory of test
files. The final key is the &lt;code&gt;commands&lt;&#x2F;code&gt; to run your test suite. Since I have the
test rule configured in my Makefile I make use of it here.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;yaml&quot; class=&quot;language-yaml &quot;&gt;&lt;code class=&quot;language-yaml&quot; data-lang=&quot;yaml&quot;&gt;test:
  imports:
    - hoomd.harmonic_force
  requires:
    - pytest
  source_files:
    - test&amp;#x2F;*
  commands:
    - make test
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;I have put all the snippets into a single block of code at the end of this article or downloadable
&lt;a href=&quot;&#x2F;code&#x2F;hoomd-plugin&#x2F;meta.yaml&quot;&gt;here&lt;&#x2F;a&gt;. For more complex examples, you can have a look at the &lt;a href=&quot;https:&#x2F;&#x2F;bitbucket.org&#x2F;glotzer&#x2F;hoomd-blue&#x2F;src&#x2F;maint&#x2F;conda-recipe&#x2F;&quot;&gt;Hoomd
repository&lt;&#x2F;a&gt; or the example in &lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;malramsay64&#x2F;hoomd-harmonic-force&#x2F;blob&#x2F;master&#x2F;meta.yaml&quot;&gt;my repository&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;building-and-distribution&quot;&gt;Building and Distribution&lt;&#x2F;h2&gt;
&lt;p&gt;With all the metadata defined, conda makes it straightforward to create the
package for upload. There are two conda packages required for building and
uploading to &lt;a href=&quot;https:&#x2F;&#x2F;anaconda.org&#x2F;&quot;&gt;Anaconda Cloud&lt;&#x2F;a&gt;, which are both required in the root
environment. You can install both packages with the command below, where the
&lt;code&gt;-n&lt;&#x2F;code&gt; flag specifies installing to the &lt;code&gt;root&lt;&#x2F;code&gt; environment.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$ conda install -n root conda-build anaconda-client
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;With the &lt;code&gt;conda-build&lt;&#x2F;code&gt; package installed, you can run &lt;code&gt;conda build .&lt;&#x2F;code&gt; where the
&lt;code&gt;.&lt;&#x2F;code&gt; is the path to the directory containing &lt;code&gt;meta.yaml&lt;&#x2F;code&gt; file. As part of build
process conda will create new environments for both the build and test phases,
preventing the packages in your current environment from interfering with the
build process and ensuring you have specified all requirements. The build phase
prepares the package for upload to &lt;a href=&quot;https:&#x2F;&#x2F;anaconda.org&#x2F;&quot;&gt;Anaconda Cloud&lt;&#x2F;a&gt;, which is handled by
anaconda client. If you have already provided credentials the upload &lt;em&gt;should&lt;&#x2F;em&gt;
be automatic, though on failure the error messages include steps to complete
the upload.&lt;&#x2F;p&gt;
&lt;p&gt;With your package uploaded to the Anaconda Cloud, anybody can install it by specifying your
repository. You can install my plugin with the command&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$ conda install -c malramsay hoomd-harmonic-force
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;although I wouldn&#x27;t recommend using it at this stage. That said, contributions
are most welcome.&lt;&#x2F;p&gt;
&lt;p&gt;While the process of packaging is difficult, I have hopefully made it somewhat
more approachable. The good thing is that once you have a &lt;code&gt;meta.yaml&lt;&#x2F;code&gt; file for
a project, maintenance will mostly be updates of version numbers. As an added
benefit, installing or updating your software is much simpler for you, and
importantly, anyone else that wants to try it out. For why open source software
if you don&#x27;t intend for someone else to actually use it.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;entire-meta-yaml&quot;&gt;Entire &lt;code&gt;meta.yaml&lt;&#x2F;code&gt;&lt;&#x2F;h2&gt;
&lt;p&gt;I have included the entire &lt;code&gt;meta.yaml&lt;&#x2F;code&gt; file below for ease of copying.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;yaml&quot; class=&quot;language-yaml &quot;&gt;&lt;code class=&quot;language-yaml&quot; data-lang=&quot;yaml&quot;&gt;# meta.yaml

package:
  name: hoomd-harmonic-force
  version: 0.1.7

about:
  home: https:&amp;#x2F;&amp;#x2F;github.com&amp;#x2F;malramsay64&amp;#x2F;hoomd-harmonic-force
  license: MIT
  license_file: LICENSE

source:
  git_url: https:&amp;#x2F;&amp;#x2F;github.com&amp;#x2F;malramsay64&amp;#x2F;hoomd-harmonic-force.git
  git_rev: v0.1.7

requirements:
  build:
    - python 3.6.*
    - setuptools
    - numpy 1.14.*
    - hoomd 2.3.*
    - cudatoolkit 8.*
    - cmake &amp;gt;=2.8.0

  run:
    - python 3.6.*
    - numpy 1.14.*
    - hoomd 2.3.*
    - cudatoolkit 8.*

build:
  script: make clean &amp;amp;&amp;amp; make install
  number: 0

test:
  imports:
    - hoomd.harmonic_force
  requires:
    - pytest
  source_files:
    - test&amp;#x2F;*
  commands:
    - make test
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>Scraping Data from Facebook</title>
        <published>2018-05-09T00:00:00+00:00</published>
        <updated>2018-05-09T00:00:00+00:00</updated>
        <author>
          <name>Unknown</name>
        </author>
        <link rel="alternate" href="https://malramsay.com/post/scraping-data-from-facebook/" type="text/html"/>
        <id>https://malramsay.com/post/scraping-data-from-facebook/</id>
        
        <content type="html">&lt;p&gt;Competition is a strange thing
making you suddenly interested in the most unusual of problems
It has become a tradition of The Lancer Band,
of which I am a member,
to produce a video as part of our ANZAC Day commemorations.
These videos have been highly successful
garnering millions of views on Facebook and
with some choice communications throughout the rest of the year
have resulted in a commendable social media following.&lt;&#x2F;p&gt;
&lt;p&gt;After a successful drive to raise the number of page likes this year,
we became more interested in how we compared to other similar pages,
from other service bands in Australia,
to other units of the Australian Army.
Rather than manually collating all this data manually,
I looked for an automated solution for this process,
both to have up to date data,
and to store historical data
allowing for comparisons with these pages over time.&lt;&#x2F;p&gt;
&lt;p&gt;The method the internet recommends for getting data from Facebook
is through their &lt;a href=&quot;https:&#x2F;&#x2F;developers.facebook.com&#x2F;docs&#x2F;graph-api&quot;&gt;GraphAPI&lt;&#x2F;a&gt;,
a method of using http requests query specific data.
The first step in using the GraphAPI is getting an access token,
the simplest method though logging into the &lt;a href=&quot;https:&#x2F;&#x2F;developers.facebook.com&#x2F;tools&#x2F;explorer&#x2F;&quot;&gt;GraphAPI Explorer&lt;&#x2F;a&gt;,
a web based tool for testing GraphAPI.
In the process of getting an access token from the GraphAPI Explorer,
a prompt will ask for the permissions granted to the access token.
You can uncheck all the permissions since they are not required for this example.
The more permissions granted to an access token,
the more data from the GraphAPI is accessible,
here interested in the &lt;code&gt;fan_count&lt;&#x2F;code&gt;, which is the number of users who like the page.
Documentation for this value,
and all the other values can are available in the &lt;a href=&quot;https:&#x2F;&#x2F;developers.facebook.com&#x2F;docs&#x2F;graph-api&#x2F;reference&#x2F;page&quot;&gt;page documentation&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;With the page and data field to query and the access token,
it is possible to query the GraphAPI in python
using the requests package as in the code example below.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;import requests
page = &amp;#x27;thelancerband&amp;#x27;
token = &amp;#x27;&amp;lt;access_token&amp;gt;&amp;#x27;
ret = requests.get(f&amp;#x27;https:&amp;#x2F;&amp;#x2F;graph.facebook.com&amp;#x2F;{page}&amp;#x27;, params={&amp;#x27;fields&amp;#x27;: &amp;#x27;fan_count&amp;#x27;, &amp;#x27;access_token&amp;#x27;: token})
ret.json()[&amp;#x27;fan_count&amp;#x27;]
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Note that you will need to replace &lt;code&gt;&amp;lt;access_token&amp;gt;&lt;&#x2F;code&gt; with the token
generated from the &lt;a href=&quot;https:&#x2F;&#x2F;developers.facebook.com&#x2F;tools&#x2F;explorer&#x2F;&quot;&gt;GraphAPI Explorer&lt;&#x2F;a&gt;.
At the time of writing &lt;a href=&quot;https:&#x2F;&#x2F;facebook.com&#x2F;thelancerband&quot;&gt;The Lancer Band&lt;&#x2F;a&gt; has 6546 likes,
which is the value output when I ran the code snippet above.
Hopefully when you are try this the result is somewhat larger.&lt;&#x2F;p&gt;
&lt;p&gt;While the GraphAPI has a relatively simple interface for querying
many different types of data,
it has a fairly significant drawback for accessing data over periods of time.
The access tokens for the GraphAPI Explorer have an expiry of 1 hour,
meaning nearly every time you want to perform data collection you need a new token.
Programmatically generating access tokens requires the creation of a Facebook application.
While this sounds simple enough,
it requires getting the app reviewed and accepted by Facebook,
far more work than I was willing to put into this.
While these requirements are fair enough for all the personal information
having access to a Facebook profile provides,
I only want to access the publicly available likes data of a page.&lt;&#x2F;p&gt;
&lt;p&gt;With all the effort required
to go through the &#x27;proper&#x27; methods to obtain this data,
I looked for alternate approaches.
As far as I can discern coming from the scientific realm,
two of the key packages for collecting and processing data from websites&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href=&quot;http:&#x2F;&#x2F;docs.python-requests.org&#x2F;en&#x2F;master&#x2F;&quot;&gt;requests&lt;&#x2F;a&gt;&lt;&#x2F;strong&gt; for an interface with http requests and responses&lt;&#x2F;li&gt;
&lt;li&gt;&lt;strong&gt;&lt;a href=&quot;https:&#x2F;&#x2F;www.crummy.com&#x2F;software&#x2F;BeautifulSoup&#x2F;bs4&#x2F;doc&#x2F;&quot;&gt;beautifulsoup4&lt;&#x2F;a&gt;&lt;&#x2F;strong&gt; which provides an interface
for extracting information from a returned html&#x2F;xml document.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;We are going to use requests to get the page which has the number of likes contained on it,
then use beautifulsoup4 to extract the number required.&lt;&#x2F;p&gt;
&lt;p&gt;Finding the web address to use in the request is a case of navigating Facebook manually,
with the number of likes appearing on the community page
having the url [https:&#x2F;&#x2F;facebook.com&#x2F;pg&#x2F;thelancerband&#x2F;community][].
The contents of this page can be downloaded using the code snippet below;&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;import requests
page_name = &amp;quot;thelancerband&amp;quot;
community_page = requests.get(f&amp;quot;https:&amp;#x2F;&amp;#x2F;facebook.com&amp;#x2F;pg&amp;#x2F;{page_name}&amp;#x2F;community&amp;quot;)
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;This gets the html,
from which we need to extract the number of likes.
Printing the entire returned page to output&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;community_page.text
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;gives a huge wall of text.
Using the search functionality to look for instances of &#x27;Total Likes&#x27;
finds the following code snippet&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;html&quot; class=&quot;language-html &quot;&gt;&lt;code class=&quot;language-html&quot; data-lang=&quot;html&quot;&gt;...&amp;lt;div class=&amp;quot;_3xom&amp;quot;&amp;gt;6,546&amp;lt;&amp;#x2F;div&amp;gt;&amp;lt;div class=&amp;quot;_3xok&amp;quot;&amp;gt;Total Likes&amp;lt;&amp;#x2F;div&amp;gt;...
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;The number of likes is in the div element preceding the div containing the text &lt;code&gt;Total Likes&lt;&#x2F;code&gt;.
Beautifulsoup4 allows us to programmatically traverse the document in a similar logic&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;from bs4 import BeautifulSoup
document = BeautifulSoup(community_page, &amp;#x27;html.parser&amp;#x27;)
document.find(text=&amp;#x27;Total Likes&amp;#x27;).parent.previous
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Here we search for text matching &#x27;Total Likes&#x27;,
which returns a reference to that location in the document.
The number of likes is the content of the previous div element,
by taking the parent element to select the surrounding div,
then the previous element we extract the value of the previous div element.&lt;&#x2F;p&gt;
&lt;p&gt;With the total number of likes as string,
we need to reformat it to convert to an integer.
This involves removing the comma as the thousands separator and
the SI prefixes &lt;code&gt;k&lt;&#x2F;code&gt;, &lt;code&gt;M&lt;&#x2F;code&gt;, etc. when the likes are in the hundreds of thousands or more.
The quick and dirty method of removing the comma is using &lt;code&gt;str.replace(&#x27;,&#x27;, &#x27;&#x27;)&lt;&#x2F;code&gt;,
though a real project should use the &lt;code&gt;locale&lt;&#x2F;code&gt; module as described on &lt;a href=&quot;https:&#x2F;&#x2F;stackoverflow.com&#x2F;questions&#x2F;1779288&#x2F;how-do-i-use-python-to-convert-a-string-to-a-number-if-it-has-commas-in-it-as-th&quot;&gt;stackoverflow&lt;&#x2F;a&gt;.
A simple fix for the prefixes is installing the &lt;a href=&quot;https:&#x2F;&#x2F;pypi.org&#x2F;project&#x2F;humanfriendly&#x2F;&quot;&gt;humanfriendly&lt;&#x2F;a&gt; package,
and using the &lt;code&gt;humanfriendly.parse_size()&lt;&#x2F;code&gt; function which will handle the prefixes.
This results in the code for getting likes from a list of pages looking like that below.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;from bs4 import BeautifulSoup
import humanfriendly
import requests

page_list = [
    &amp;quot;thelancerband&amp;quot;,
]

def get_num_likes(page_id):
    community_page = requests.get(f&amp;quot;https:&amp;#x2F;&amp;#x2F;facebook.com&amp;#x2F;pg&amp;#x2F;{page_id}&amp;#x2F;community&amp;quot;)
    document = BeautifulSoup(community_page, &amp;#x27;html.parser&amp;#x27;)
    page_likes = humanfriendly.parse_size(
        document.find(text=&amp;#x27;Total Likes&amp;#x27;).parent.previous.replace(&amp;#x27;,&amp;#x27;, &amp;#x27;&amp;#x27;)
    )
    return page_likes

for page in page_list:
    print(f&amp;#x27;The Facebook page for {page} has {get_num_likes(page)} likes&amp;#x27;)
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;This prints the results to the screen,
which is not useful for longer term analysis
although it is easy to see that everything is working as expected.&lt;&#x2F;p&gt;
&lt;p&gt;There are many different methods for storing this data for later analysis,
from a simple csv file to a sqlite database.
I am most familiar with pandas and HDF5 files,
hence these were my tools of choice
resulting in the data collection loop below.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;import pandas
data = []
for page in page_list:
    data.append({
        &amp;#x27;page_id&amp;#x27;: page,
        &amp;#x27;likes&amp;#x27;: get_num_likes(page),
        &amp;#x27;time&amp;#x27;: pandas.Timestamp.now(),
    })
df = pandas.DataFraame.from_records(data)
df.set_index(&amp;#x27;time&amp;#x27;).to_hdf(&amp;#x27;likes_data.h5&amp;#x27;, &amp;#x27;data&amp;#x27;, format=&amp;#x27;table&amp;#x27;, append=True)
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Note that the &lt;code&gt;tables&lt;&#x2F;code&gt; package is also a requirement for this code snippet to work.&lt;&#x2F;p&gt;
&lt;p&gt;This is my first foray into scraping websites for data analysis
which is surprisingly simple with the appropriate tools.
The biggest issue I encountered in this project
was trying to understand the Facebook authentication tokens.
While this is a simple example,
it can easily be extended for different use cases.&lt;&#x2F;p&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>A guide to setting up remote SSH</title>
        <published>2018-04-09T00:00:00+00:00</published>
        <updated>2018-04-09T00:00:00+00:00</updated>
        <author>
          <name>Unknown</name>
        </author>
        <link rel="alternate" href="https://malramsay.com/post/remote-ssh-configuration/" type="text/html"/>
        <id>https://malramsay.com/post/remote-ssh-configuration/</id>
        
        <content type="html">&lt;p&gt;In the right (or wrong) hands ssh is a powerful tool
for the remote management of a Unix system.
Most desktop, or workstation distributions of Linux disable
remote access over ssh by default.
The simplest method to check if you have ssh server running on your machine
is to run&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$ ssh localhost
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;If ssh is not installed or running
this will print out a message&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code&gt;ssh: connect to host localhost port 22: Connection refused
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;most likely indicating that the ssh server is not running.&lt;&#x2F;p&gt;
&lt;p&gt;Before heading any further in enabling remote access over ssh,
it is a good idea to ensure that you have a strong password
for all of the accounts on the device,
in particular the root account (even better is disabling it).
By enabling remote access for yourself
you are also enabling remote access for anyone else that might want to try and log in.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;installation&quot;&gt;Installation&lt;&#x2F;h2&gt;
&lt;p&gt;The package required for installation is in most distributions named &lt;code&gt;openssh-server&lt;&#x2F;code&gt;.
So for Ubuntu running the command;&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$ sudo apt install openssh-server
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;will install the package as appropriate.
Other Linux distributions have different package managers so
use the one for your distribution.
So for fedora, replace &lt;code&gt;apt&lt;&#x2F;code&gt; with &lt;code&gt;dnf&lt;&#x2F;code&gt; and for CentOS replace &lt;code&gt;apt&lt;&#x2F;code&gt; with &lt;code&gt;yum&lt;&#x2F;code&gt;.
On macOS everything is already installed and just needs to be enabled
in System Preferences &amp;gt; Sharing then ensure Remote Login is checked.&lt;&#x2F;p&gt;
&lt;p&gt;Once the package is installed,
it needs to be both added to the list of packages to run at boot,
and started now to test.&lt;&#x2F;p&gt;
&lt;p&gt;To start the openssh server, run the below command&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$ sudo systemctl start ssh
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;which is using the &lt;code&gt;systemd&lt;&#x2F;code&gt; init system to start the openssh server instance.
For fedora and CentOS, the ssh service instead has the name sshd so run&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$ sudo systemctl start sshd
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;to start the server instance.
Since this is probably a service you want running automatically on boot,
running the command&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$ sudo systemctl enable ssh
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;will put files in the appropriate places to enable the service on boot.
Other useful commands for &lt;code&gt;systemctl&lt;&#x2F;code&gt; are &lt;code&gt;stop&lt;&#x2F;code&gt;, &lt;code&gt;restart&lt;&#x2F;code&gt;, and &lt;code&gt;disable&lt;&#x2F;code&gt;&lt;&#x2F;p&gt;
&lt;p&gt;Once the ssh server is running
our command to connect at localhost&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$ ssh localhost
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;should present a new message like the one below.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;txt&quot; class=&quot;language-txt &quot;&gt;&lt;code class=&quot;language-txt&quot; data-lang=&quot;txt&quot;&gt;The authenticity of host &amp;#x27;localhost (::1)&amp;#x27; can&amp;#x27;t be established.
ECDSA key fingerprint is ba:ba:b1:4c:5a:e2:aa:40:aa:7f:aa:01:aa:df:aa:47.
Are you sure you want to continue connecting (yes&amp;#x2F;no)?
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;In some ways this is like the terms and conditions when you sign up for a website
in that you should care about and understand it,
however, typically you click through (or type yes) without reading.
Since we are connecting to localhost,
which has an IPV6 address of &lt;code&gt;::1&lt;&#x2F;code&gt;
or IPV4 address of &lt;code&gt;127.0.0.1&lt;&#x2F;code&gt;
we can trust the key fingerprint.
This is a method of checking that whenever we log into a host over ssh,
the host hasn&#x27;t changed which if unexpected
could be indicative of a man-in-the-middle attack or other nefarious actions.&lt;&#x2F;p&gt;
&lt;p&gt;Now we are able to connect to the machine locally,
we need to work out how to connect from a remote machine.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;how-do-i-find-my-device&quot;&gt;How do I Find My Device?&lt;&#x2F;h2&gt;
&lt;p&gt;With the previous commands we have been passing &lt;code&gt;localhost&lt;&#x2F;code&gt; to the ssh command,
in the same way we might type &lt;code&gt;google.com&lt;&#x2F;code&gt; into a web browser.
These names are a human readable format for addressing devices,
which map to IPV6 or IPV4 addresses computers read to initiate a connection.&lt;&#x2F;p&gt;
&lt;p&gt;To find the name of the computer we are working on,
also known as the hostname,
we can use the &lt;code&gt;hostname&lt;&#x2F;code&gt; command.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$ hostname -f
lovelace.staff.sydney.edu.au
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Here the &lt;code&gt;-f&lt;&#x2F;code&gt; gives the full hostname of the device you are using,
with a &lt;code&gt;-s&lt;&#x2F;code&gt; flag also supported,
however we need to full hostname.&lt;&#x2F;p&gt;
&lt;p&gt;With the full hostname,
we know the name of our device,
and we know that our device is aware of it&#x27;s hostname.
This doesn&#x27;t necessarily mean that any other devices know what ours is called.
The process of taking a hostname and turning it into an address
the computer can use is known as DNS,
which is provided by name servers.
Typically in organisations there are multiple levels of name servers,
which will handle both internal names,
like the hostname of our device,
and external names
like &lt;code&gt;github.com&lt;&#x2F;code&gt; or &lt;code&gt;wikipedia.org&lt;&#x2F;code&gt;.
Where a device is querying an external name server,
it will not be able to resolve local names.
To test this we can use the name server lookup &lt;code&gt;nslookup&lt;&#x2F;code&gt; command.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;txt&quot; class=&quot;language-txt &quot;&gt;&lt;code class=&quot;language-txt&quot; data-lang=&quot;txt&quot;&gt;$ nslookup lovelace.staff.sydney.edu.au
Server:         172.18.240.210
Address:        172.18.240.210#53

Name:   lovelace.staff.sydney.edu.au
Address: 10.65.205.200
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;This tells us that the server at the ip address &lt;code&gt;172.18.240.210&lt;&#x2F;code&gt;
which my laptop is using as the default name server,
knows that &lt;code&gt;lovelace.staff.sydney.edu.au&lt;&#x2F;code&gt; exists at the address &lt;code&gt;10.65.205.200&lt;&#x2F;code&gt;.
If we instead used one of the large public DNS servers
provided by Google (&lt;code&gt;8.8.8.8&lt;&#x2F;code&gt;) or CloudFlare (&lt;code&gt;1.1.1.1&lt;&#x2F;code&gt;)
we will instead get an error message.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;txt&quot; class=&quot;language-txt &quot;&gt;&lt;code class=&quot;language-txt&quot; data-lang=&quot;txt&quot;&gt;$ nslookup lovelace.staff.sydney.edu.au 1.1.1.1
Server:         1.1.1.1
Address:        1.1.1.1#53

** server can&amp;#x27;t find lovelace.staff.sydney.edu.au: SERVFAIL
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;h2 id=&quot;setting-the-hostname&quot;&gt;Setting the hostname&lt;&#x2F;h2&gt;
&lt;p&gt;When going through the initial setup of macOS or linux machines,
there is a step where you give the machine the hostname.
To later change the hostname requires the editing a couple of files.
The file &lt;code&gt;&#x2F;etc&#x2F;hostname&lt;&#x2F;code&gt; stores the hostname of the machine,
which will typically be the &lt;a href=&quot;https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Fully_qualified_domain_name&quot;&gt;Fully Qualified Domain Name&lt;&#x2F;a&gt; (FQDN) of the host.
To actually change the hostname a reboot is required,
as this file is read during the boot sequence.
Before you do reboot,
there is another file &lt;code&gt;&#x2F;etc&#x2F;hosts&lt;&#x2F;code&gt; that requires editing.&lt;&#x2F;p&gt;
&lt;p&gt;The &lt;code&gt;&#x2F;etc&#x2F;hosts&lt;&#x2F;code&gt; file is the fist place that is checked
when looking to resolve a domain name
and is sometimes used as a way to block websites
or just to resolve hostname.
We want to our domain name to resolve to &lt;code&gt;127.0.0.1&lt;&#x2F;code&gt; and &lt;code&gt;::1&lt;&#x2F;code&gt;
so that services like web servers don&#x27;t get confused.
Ensure the following lines (with the appropriate hostname) are in the &lt;code&gt;&#x2F;etc&#x2F;hosts&lt;&#x2F;code&gt; file&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;txt&quot; class=&quot;language-txt &quot;&gt;&lt;code class=&quot;language-txt&quot; data-lang=&quot;txt&quot;&gt;127.0.0.1   lovelace lovelace.malramsay.com
::1         lovelace lovelace.malramsay.com
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;This sets the ip address we want the hostname to resolve to,
which on the first line is &lt;code&gt;127.0.0.1&lt;&#x2F;code&gt;,
followed by a list of hostnames,
being the short version &lt;code&gt;lovelace&lt;&#x2F;code&gt; and the FQDN &lt;code&gt;lovelace.malramsay.com&lt;&#x2F;code&gt;.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;what-do-i-do-if-i-can-t-lookup-my-hostname&quot;&gt;What do I do if I can&#x27;t lookup my hostname?&lt;&#x2F;h2&gt;
&lt;p&gt;The reason we use domain names and hostnames
is that they are far easier to remember than a string of numbers.
It is possible to navigate to Google by navigating to the ip address &lt;a href=&quot;http:&#x2F;&#x2F;216.58.200.110&quot;&gt;216.58.200.110&lt;&#x2F;a&gt;
which is the address returned for google.com.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;txt&quot; class=&quot;language-txt &quot;&gt;&lt;code class=&quot;language-txt&quot; data-lang=&quot;txt&quot;&gt;$ nslookup google.com
Server:         172.18.240.210
Address:        172.18.240.210#53

Non-authoritative answer:
Name:   google.com
Address: 216.58.200.110
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;While using the IP address is possible,
it is far more cumbersome to remember and to type.
For a single IP address which we can put in a configuration file,
this isn&#x27;t too problematic.&lt;&#x2F;p&gt;
&lt;p&gt;To find the ip address of the current host
run the command &lt;code&gt;ip addr&lt;&#x2F;code&gt;,
which will give output like that below.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;txt&quot; class=&quot;language-txt &quot;&gt;&lt;code class=&quot;language-txt&quot; data-lang=&quot;txt&quot;&gt;$ ip addr
lo0: flags=8049&amp;lt;UP,LOOPBACK,RUNNING,MULTICAST&amp;gt; mtu 16384
        inet 127.0.0.1&amp;#x2F;8 lo0
        inet6 ::1&amp;#x2F;128
        inet6 fe80::1&amp;#x2F;64 scopeid 0x1
en0: flags=8863&amp;lt;UP,BROADCAST,SMART,RUNNING,SIMPLEX,MULTICAST&amp;gt; mtu 1500
        ether 78:4f:43:4d:cc:53
        inet6 fe80::18c5:93d4:e972:b798&amp;#x2F;64 secured scopeid 0x7
 &amp;gt;      inet 10.16.249.156&amp;#x2F;21 brd 10.16.255.255 en0
awdl0: flags=8943&amp;lt;UP,BROADCAST,RUNNING,PROMISC,SIMPLEX,MULTICAST&amp;gt; mtu 1484
        ether b6:54:4e:bd:a7:d7
        inet6 fe80::b454:4eff:febd:a7d7&amp;#x2F;64 scopeid 0xa
utun0: flags=8051&amp;lt;UP,POINTOPOINT,RUNNING,MULTICAST&amp;gt; mtu 2000
        inet6 fe80::1c95:35f4:acea:1034&amp;#x2F;64 scopeid 0x10
en5: flags=8863&amp;lt;UP,BROADCAST,SMART,RUNNING,SIMPLEX,MULTICAST&amp;gt; mtu 1500
        ether ac:df:48:00:11:22
        inet6 fe80::aede:48ff:fe00:1122&amp;#x2F;64 scopeid 0x9
en7: flags=8863&amp;lt;UP,BROADCAST,SMART,RUNNING,SIMPLEX,MULTICAST&amp;gt; mtu 1500
        ether 00:0e:c6:e2:bb:69
        inet6 fe80::1871:5dc7:1f58:2127&amp;#x2F;64 secured scopeid 0x12
 &amp;gt;      inet 10.65.205.200&amp;#x2F;24 brd 10.65.205.255 en7
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;The lines we are looking for are those with &lt;code&gt;inet&lt;&#x2F;code&gt; at the start
which I have denoted with a &lt;code&gt;&amp;gt;&lt;&#x2F;code&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;This gives me two IP addresses,
one for the wifi and another for the wired ethernet connection.
You may have noticed that the &lt;code&gt;10.65.205.200&lt;&#x2F;code&gt; address matched that obtained
from the &lt;code&gt;nslookup&lt;&#x2F;code&gt; of my hostname above.
This tells me that the name server is directing
queries of my hostname back to my device.&lt;&#x2F;p&gt;
&lt;p&gt;Now we have an ip address,
how do we make the computer remember it
when the name server doesn&#x27;t.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;remembering-ip-addresses&quot;&gt;Remembering IP addresses&lt;&#x2F;h2&gt;
&lt;p&gt;The DNS lookup of IP addresses from hostnames works on
the institutional network I have access to at the University of Sydney.
You will likely not have this same configuration,
so there are some alternate methods
for getting your computer to remember ip addresses.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;ssh-configuration-file&quot;&gt;SSH configuration file&lt;&#x2F;h3&gt;
&lt;p&gt;The simplest method of getting the computer
to remember an ip address for you is the ssh config file &lt;code&gt;~&#x2F;.ssh&#x2F;config&lt;&#x2F;code&gt;.
This approach only works when connecting using ssh,
however, it is useful for setting all the parameters for an ssh connection.
By putting the code snippet below in &lt;code&gt;~&#x2F;.ssh&#x2F;config&lt;&#x2F;code&gt;&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sshconfig&quot; class=&quot;language-sshconfig &quot;&gt;&lt;code class=&quot;language-sshconfig&quot; data-lang=&quot;sshconfig&quot;&gt;Host lovelace
    Hostname 10.65.205.200
    User malcolm
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;running the command &lt;code&gt;ssh lovelace&lt;&#x2F;code&gt;,
is equivalent to running &lt;code&gt;ssh malcolm@10.65.205.200&lt;&#x2F;code&gt;
and is much easier to both type and remember (and tab-complete).
Additional useful parameters for the configuration are&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;LocalForward&lt;&#x2F;code&gt;&#x2F;&lt;code&gt;RemoteForward&lt;&#x2F;code&gt; for port forwarding,&lt;&#x2F;li&gt;
&lt;li&gt;&lt;code&gt;Port&lt;&#x2F;code&gt; when using a non-standard port, and&lt;&#x2F;li&gt;
&lt;li&gt;&lt;code&gt;IdentityFile&lt;&#x2F;code&gt; for using specific private keys.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;The &lt;code&gt;ssh_config&lt;&#x2F;code&gt; man pages (&lt;code&gt;man ssh_config&lt;&#x2F;code&gt;) also contain
all the parameters available and information about each of them.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;hosts-file&quot;&gt;Hosts file&lt;&#x2F;h3&gt;
&lt;p&gt;Earlier we modified the &lt;code&gt;&#x2F;etc&#x2F;hosts&lt;&#x2F;code&gt; file to change the hostname of our system.
We can also use it to remember the ip address of other systems.
This is not recommended for managing many devices, use DNS for that,
for a small number of systems we can use use this file for the translation of
domain names to ip addresses.
Instead of using the ip address of localhost,
we use the ip address we want the hostname to resolve to as shown below.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;txt&quot; class=&quot;language-txt &quot;&gt;&lt;code class=&quot;language-txt&quot; data-lang=&quot;txt&quot;&gt;10.65.205.200   lovelace.malramsay.com
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;The advantage this approach has over the ssh configuration
is that it will also resolve the hostname in the web browser.
This also makes it much easier to access web services
like jupyter notebooks on the remote host.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;conclusion&quot;&gt;Conclusion&lt;&#x2F;h2&gt;
&lt;p&gt;This should be sufficient to get started using SSH within the same network.
To access the service over the internet,
where possible use the institutional VPN,
or set up your own.&lt;&#x2F;p&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>Adventures in the Python Visualisation Landscape</title>
        <published>2018-03-17T00:00:00+00:00</published>
        <updated>2018-03-17T00:00:00+00:00</updated>
        <author>
          <name>Unknown</name>
        </author>
        <link rel="alternate" href="https://malramsay.com/post/adventures-in-visualisation/" type="text/html"/>
        <id>https://malramsay.com/post/adventures-in-visualisation/</id>
        
        <content type="html">&lt;p&gt;Calling the current python visualisation landscape fragmented would probably be an understatement,
since itself requires a &lt;a href=&quot;https:&#x2F;&#x2F;youtu.be&#x2F;FytuB8nFHPQ?t=3m53s&quot;&gt;visualisation&lt;&#x2F;a&gt; to even begin to comprehend.
For various reasons I have been unhappy with the tool I was using for visualisation at that time
and I have been searching for the one visualisation package to rule them all (spoiler: it doesn&#x27;t exist)&lt;&#x2F;p&gt;
&lt;p&gt;About a year ago I was using Matplotlib for all my figures,
which enabled me to create anything I wanted,
usually with a stackoverflow answer giving me a working example to adapt.
Problem was, nearly everything required a stackoverflow post or the documentation,
making the process of building these figures slow and tedious.
Around the same time I was getting fed up with Matplotlib,
I found Jake VanderPlas&#x27; PyCon talk on &lt;a href=&quot;https:&#x2F;&#x2F;youtu.be&#x2F;FytuB8nFHPQ?t=3m53s&quot;&gt;The Python Visualisation Landscape&lt;&#x2F;a&gt;
which has since acted as both a map and a list of achievements to unlock.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;altair-1-2&quot;&gt;Altair 1.2&lt;&#x2F;h3&gt;
&lt;p&gt;Somewhat naturally the first package I looked at for improving my visualisation workflow was
Altair which was at the time a 1.2 release.
At the time I thought it was fine for simple figures,
however my data was not formatted to make the most of what Altair could offer,
which as a relative newbie to pandas was frustrating.
Another sticking point with Altair was working out how to customise figures,
After Matplotlib, I was expecting to search what I wanted to in Google
and have appropriate stack overflow answer on the first page of results.
With Altair being a new package, this how-to style documentation didn&#x27;t exist,
which meant I never worked out how to customise the figure.
The most frustrating of the default settings was the axis labels
which use the SI Prefix for large numbers
i.e. 1Mm for $1 \times 10^6$  and 1no for $1 \times 10^{-9}$,
something which I never worked out how to change at the time.
Another problem I had with Altair (and Matplotlib) was the lack of interactivity.
I was in the early stages of data investigation
which required both a high level overview of the trends,
while also being able to look at smaller regions in more detail.
Having to constantly change axis ranges was
making it slow and frustrating to create a figure with the appropriate view.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;bokeh&quot;&gt;Bokeh&lt;&#x2F;h3&gt;
&lt;p&gt;With an emphasis on interactivity,
&lt;a href=&quot;https:&#x2F;&#x2F;bokeh.pydata.org&#x2F;en&#x2F;latest&#x2F;&quot;&gt;Bokeh&lt;&#x2F;a&gt; was the tool I was looking for to investigate data on different scales.
With the ability to generate figures in a notebook,
for quick visualisations of datasets to understand the data,
and compare it with previous studies,
and also as a &lt;a href=&quot;https:&#x2F;&#x2F;bokeh.pydata.org&#x2F;en&#x2F;latest&#x2F;docs&#x2F;user_guide&#x2F;server.html&quot;&gt;Bokeh server&lt;&#x2F;a&gt; application,
to perform standard analyses on large volumes of data,
like ensuring a simulation is running properly.
Despite having these interactive visualisations,
bokeh is lacking in the same way as Matplotlib,
it takes a long time to specify everything required to create the figures.
While there was a start on a high level plotting interface in the form of &lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;bokeh&#x2F;bkcharts&quot;&gt;bkcharts&lt;&#x2F;a&gt;,
the code is now unmaintained and directs users to Holoviews.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;holoviews&quot;&gt;Holoviews&lt;&#x2F;h3&gt;
&lt;p&gt;Like Altair, &lt;a href=&quot;https:&#x2F;&#x2F;holoviews.org&quot;&gt;Holoviews&lt;&#x2F;a&gt; is a declarative plotting interface,
providing a quick and simple interface for constructing figures.
Holoviews is structures around the idea of 
describing data on creation of the dataset,
rather than the construction of the figure,
allowing for simple figure definitions.
Rather than being a library that actually creates a figure,
Holoviews performs the reasoning about the dataset
and passing the rendering off to Bokeh, Matplotlib, or Plotly.
I initially saw this as a big strength of Holoviews,
that I could generate interactive visualisations in bokeh,
change the output to Matplotlib and have a configurable figure for publication.
In practice this is not so simple,
with some modifications to plot style parameters
changing between the different output formats
and having a limited range of customisations.&lt;&#x2F;p&gt;
&lt;p&gt;The drawback of having this special annotated data object
is that you lose all the flexibility of having a pandas DataFrame
and the vast array of operations that allows.
This is important to me 
because my field of science has a long history of 
researchers taking some quantities and 
combining them to give a describable temperature dependence.
This lack of flexibility led me back to Altair,
this time version 2.0.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;altair-2-0&quot;&gt;Altair 2.0&lt;&#x2F;h3&gt;
&lt;p&gt;Why Altair again?
When I first tried Altair I was approaching it from Matplotlib,
with its extensive documentation
in the form of the technical reference and the numerous how to guides.
This time I was approaching it from Bokeh and Holoviews,
both of which are relatively new libraries with some teething problems.
As a consequence I had become much better at problem solving issues
and navigating technical reference materials.
Another major moment of discovery was making the connection 
that since Altair implements the Vega-Lite specification,
I should have a look at the &lt;a href=&quot;https:&#x2F;&#x2F;vega.github.io&#x2F;vega-lite&#x2F;docs&#x2F;&quot;&gt;Vega-Lite documentation&lt;&#x2F;a&gt;.
This turned out to be particularly helpful 
because the Vega-Lite documentation 
is currently more extensive than for &lt;a href=&quot;https:&#x2F;&#x2F;altair-viz.github.io&#x2F;index.html&quot;&gt;Altair&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;It was in reading the Vega-Lite documentation 
that I finally understood how to use transform functions in Altair.
These are a set of functions that perform computations on the input dataset to generate the resulting figure.
This allows me to have a single canonical dataset,
with data transformations like ratios of two quantities tied to the figure,
rather than following awkwardly named variables around.
One example of these functions is 
to make the data in the cars dataset 
useful for the majority of the world&#x27;s population
by converting the units as part of the figure definition.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;import altair as alt
from vega_datasets import data

cars = data.cars()

chart = alt.Chart(cars).mark_circle().transform_filter(
    alt.expr.datum.Miles_per_Gallon &amp;gt; 0
).transform_calculate(
    &amp;#x27;Fuel Economy (L&amp;#x2F;100 km)&amp;#x27;, &amp;#x27;235.2 &amp;#x2F; datum.Miles_per_Gallon&amp;#x27;
).transform_calculate(
    &amp;#x27;Weight (kg)&amp;#x27;, &amp;#x27;datum.Weight_in_lbs * 0.45&amp;#x27;
)
chart.encode(
    x=&amp;#x27;Fuel Economy (L&amp;#x2F;100 km):Q&amp;#x27;,
    y=&amp;#x27;Weight (kg):Q&amp;#x27;,
    size=&amp;#x27;Acceleration:Q&amp;#x27;
)
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;&lt;img src=&quot;&#x2F;img&#x2F;altair-cars-metric.svg&quot; alt=&quot;Fuel economy (L&#x2F;100km) vs weight (kg) from the cars dataset.&quot; &#x2F;&gt;&lt;&#x2F;p&gt;
&lt;p&gt;This allows for only storing the fundamental values in the dataset
and being able to compute derived values as part of the figure,
something that is particularly useful in my workflow.&lt;&#x2F;p&gt;
&lt;p&gt;This computing of values,
also extends to the computation of histograms,
complete with shortened notation.
Using the chart object from above it is possible to easily create a histogram&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;chart.mark_bar().encode(
    x=alt.X(&amp;#x27;Fuel Economy (L&amp;#x2F;100 km):Q&amp;#x27;, bin=True),
    y=&amp;#x27;count():Q&amp;#x27;,
)
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;&lt;img src=&quot;&#x2F;img&#x2F;altair-cars-hist.svg&quot; alt=&quot;Histogram of the fuel economy in the cars dataset.&quot; &#x2F;&gt;&lt;&#x2F;p&gt;
&lt;p&gt;Where setting &lt;code&gt;bin=True&lt;&#x2F;code&gt; will create bins with the default parameters,
and the &lt;code&gt;count():Q&lt;&#x2F;code&gt; on the &lt;code&gt;y&lt;&#x2F;code&gt; axis counts the elements in each bin.
Instead of &lt;code&gt;count()&lt;&#x2F;code&gt; it is also possible to perform other &lt;a href=&quot;https:&#x2F;&#x2F;vega.github.io&#x2F;vega-lite&#x2F;docs&#x2F;aggregate.html#ops&quot;&gt;aggregations&lt;&#x2F;a&gt;,
like computing the mean of a column.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;chart.mark_bar().encode(
    x=alt.X(&amp;#x27;Fuel Economy (L&amp;#x2F;100 km):Q&amp;#x27;, bin=True),
    y=&amp;#x27;mean(Weight (kg)):Q&amp;#x27;,
)
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;&lt;img src=&quot;&#x2F;img&#x2F;altair-cars-weight.svg&quot; alt=&quot;Histogram of the fuel economy in the cars dataset.&quot; &#x2F;&gt;&lt;&#x2F;p&gt;
&lt;p&gt;For a more comprehensive view of using Altiar,
have a look at either the &lt;a href=&quot;https:&#x2F;&#x2F;altair-viz.github.io&#x2F;gallery&#x2F;index.html&quot;&gt;Example Gallery&lt;&#x2F;a&gt;,
or a &lt;a href=&quot;https:&#x2F;&#x2F;altair-viz.github.io&#x2F;case_studies&#x2F;exploring-weather.html&quot;&gt;case study&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
&lt;p&gt;Each of the visualisation libraries in python
have their own strengths and weaknesses,
types of visualisations they excel at,
and others you wouldn&#x27;t want to try.
For me, while Altair does still have a some quirks,
most notably in the handling of &lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;altair-viz&#x2F;altair&#x2F;issues&#x2F;249&quot;&gt;large datasets&lt;&#x2F;a&gt;,
and a somewhat complicated method of &lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;altair-viz&#x2F;altair&#x2F;issues&#x2F;585&quot;&gt;setting titles&lt;&#x2F;a&gt;,
it provides a simple and intuitive interface to data
which at the time of writing makes it the first tool I will reach for
to understand a dataset.&lt;&#x2F;p&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>Creating an HDF5 Bomb</title>
        <published>2018-03-15T00:00:00+00:00</published>
        <updated>2018-03-15T00:00:00+00:00</updated>
        <author>
          <name>Unknown</name>
        </author>
        <link rel="alternate" href="https://malramsay.com/post/creating-hdf5-bomb/" type="text/html"/>
        <id>https://malramsay.com/post/creating-hdf5-bomb/</id>
        
        <content type="html">&lt;p&gt;You may have heard of a &lt;a href=&quot;https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;Zip_bomb&quot;&gt;zip bomb&lt;&#x2F;a&gt; or other decompression &#x27;bombs&#x27;,
which have the basic premise of containing a large volume of highly redundant data
that when decompressed takes up more resources than the system can handle.
Within the HDF5 file format there is &lt;a href=&quot;https:&#x2F;&#x2F;support.hdfgroup.org&#x2F;HDF5&#x2F;faq&#x2F;compression.html&quot;&gt;support for compression&lt;&#x2F;a&gt;,
an excellent tool for reducing file sizes,
however also ripe for exploitation.
This &#x27;issue&#x27;&lt;sup class=&quot;footnote-reference&quot;&gt;&lt;a href=&quot;#1&quot;&gt;1&lt;&#x2F;a&gt;&lt;&#x2F;sup&gt; of a file containing far more data expected, whether accidental or malicious,
is not limited to HDF5 files, any filetype supporting compression is susceptible.&lt;&#x2F;p&gt;
&lt;p&gt;How do you actually create one a decompression bomb.
The simplest method in python is to create a pandas DataFrame
comprising lots of strings which are all the same.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;import pandas as pd
large_number = 1_000_000
df = pd.DataFrame({&amp;#x27;evil&amp;#x27;: [&amp;#x27;😈&amp;#x27;]*large_number})
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Using the &lt;code&gt;sys.getsizeof&lt;&#x2F;code&gt; function we can find
the size in memory of this DataFrame&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;import sys.getsizeof
sys.getsizeof(df) &amp;#x2F; 1024 &amp;#x2F; 1024  # bytes &amp;#x2F; kilobytes &amp;#x2F; megabytes
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;which turns out to be 840 MB,
10 times larger than a DataFrame with the same number of integers.
This significant overhead of using strings is because
each element in the DataFrame contains all the storage overhead of a python object,&lt;sup class=&quot;footnote-reference&quot;&gt;&lt;a href=&quot;#2&quot;&gt;2&lt;&#x2F;a&gt;&lt;&#x2F;sup&gt;
rather than each column for numerical types,
and is the main reason I chose strings for this diabolical construct.&lt;&#x2F;p&gt;
&lt;p&gt;Now we have a large, low entropy dataset we need to save it to disk.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;with pd.HDFStore(&amp;#x27;dataset.h5&amp;#x27;, mode=&amp;#x27;w&amp;#x27;, complevel=9, complib=&amp;#x27;bzip2&amp;#x27;) as dst:
    dst.append(&amp;#x27;data&amp;#x27;, df)
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;A number of keyword arguments are set in opening the HDFStore file handle,&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;mode=&#x27;w&#x27;&lt;&#x2F;code&gt; ensures that a file with the same name is overwritten if it exists, making testing this
out much simpler.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;code&gt;complevel=9&lt;&#x2F;code&gt; sets the compression to the maximum possible not caring about the processing
requirements of doing so.&lt;&#x2F;li&gt;
&lt;li&gt;&lt;code&gt;complib=&#x27;bzip2&#x27;&lt;&#x2F;code&gt; sets the compression library to &lt;code&gt;bzip2&lt;&#x2F;code&gt; which had significantly better
compression ratios than the default &lt;code&gt;zlib&lt;&#x2F;code&gt; in my testing of this.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;Additionally the &lt;code&gt;append&lt;&#x2F;code&gt; function has been used to write the DataFrame to disk
since we want to create a dataframe that doesn&#x27;t fit in memory,
which will require appending to the file numerous times.
With these options, the 840 MB DataFrame is a 2.2 MB HDF5 file on disk,
a compression ratio greater than 200.&lt;&#x2F;p&gt;
&lt;p&gt;With all the pieces in place we can now construct our file.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;num_iters = 20
with pd.HDFStore(&amp;#x27;dataset.h5&amp;#x27;, mode=&amp;#x27;w&amp;#x27;, complevel=9, complib=&amp;#x27;bzip2&amp;#x27;) as dst:
    for _ in range(num_iters):
        dst.append(&amp;#x27;data&amp;#x27;, df)
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;This generates a &lt;a href=&quot;https:&#x2F;&#x2F;drive.google.com&#x2F;open?id=1tlr00OFEuKMkSz0slczInh3211-ZRBIz&quot;&gt;small unsuspecting 42 MB file&lt;&#x2F;a&gt;&lt;sup class=&quot;footnote-reference&quot;&gt;&lt;a href=&quot;#3&quot;&gt;3&lt;&#x2F;a&gt;&lt;&#x2F;sup&gt;
which when loaded as a pandas DataFrame becomes an 18 GB object in memory,
a compression ratio of 440.&lt;&#x2F;p&gt;
&lt;p&gt;While this is a little fun and devious,
it does highlight the importance of thinking about how we represent the data we are processing,
in particular text data.
I originally came across this HDF5 bomb by accident,
leaving a field as text when it should have been a category.
In this devious case presented, using the type &#x27;category&#x27; in the DataFrame&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code&gt;df[&amp;#x27;evil&amp;#x27;] = df[&amp;#x27;evil&amp;#x27;].astype(&amp;#x27;category&amp;#x27;)
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;our 840 MB DataFrame becomes 10 MB,
and there are no issues storing 20 of them in memory.
While it may be cool to use big data tools like Spark, Dask, or Hadoop
sometimes the simplest approach is to make the big data small.&lt;&#x2F;p&gt;
&lt;div class=&quot;footnote-definition&quot; id=&quot;1&quot;&gt;&lt;sup class=&quot;footnote-definition-label&quot;&gt;1&lt;&#x2F;sup&gt;
&lt;p&gt;The compression is working exactly as intended, it just hides the true size of the underlying data.&lt;&#x2F;p&gt;
&lt;&#x2F;div&gt;
&lt;div class=&quot;footnote-definition&quot; id=&quot;2&quot;&gt;&lt;sup class=&quot;footnote-definition-label&quot;&gt;2&lt;&#x2F;sup&gt;
&lt;p&gt;From interrogating the size of the DataFrame using either &lt;code&gt;sys.getsizeof(df)&lt;&#x2F;code&gt; or &lt;code&gt;df.memeory_useage(deep=True)&lt;&#x2F;code&gt; it appears that the memory is allocated for each object. When querying the individual objects using &lt;code&gt;id&lt;&#x2F;code&gt;, they all return the same value, which is the same as just the string. I don&#x27;t know what is going on and would be happy for someone to point me to a good resource. &lt;&#x2F;p&gt;
&lt;&#x2F;div&gt;
&lt;div class=&quot;footnote-definition&quot; id=&quot;3&quot;&gt;&lt;sup class=&quot;footnote-definition-label&quot;&gt;3&lt;&#x2F;sup&gt;
&lt;p&gt;Of course you are going to download a file some random stranger on the internet tells you is going to crash your python interpreter.&lt;&#x2F;p&gt;
&lt;&#x2F;div&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>The Perils of Packaging in Python</title>
        <published>2018-01-07T00:00:00+00:00</published>
        <updated>2018-01-07T00:00:00+00:00</updated>
        <author>
          <name>Unknown</name>
        </author>
        <link rel="alternate" href="https://malramsay.com/post/perils-of-packaging/" type="text/html"/>
        <id>https://malramsay.com/post/perils-of-packaging/</id>
        
        <content type="html">&lt;p&gt;There are many guides to packaging a python application,
including the official &lt;a href=&quot;https:&#x2F;&#x2F;packaging.python.org&#x2F;&quot;&gt;Python Packaging User Guide&lt;&#x2F;a&gt;.
While these guides offer step by step instructions for
deploying a simple application,
when deviating from a guide there can be unexpected problems.&lt;&#x2F;p&gt;
&lt;p&gt;The application I have been trying to deploy is &lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;malramsay64&#x2F;statdyn-analysis&quot;&gt;this one&lt;&#x2F;a&gt;,
a collection of tools to assist in the analysis of
molecular dynamics trajectories for my PhD.
Most of the code I have written is python,
with small sections of Cython code for performance.
I have been running automated testing on Travis-CI
and have made many different attempts at deploying the package
to both &lt;a href=&quot;https:&#x2F;&#x2F;anaconda.org&#x2F;&quot;&gt;Anaconda Cloud&lt;&#x2F;a&gt; and &lt;a href=&quot;https:&#x2F;&#x2F;pypi.python.org&#x2F;pypi&quot;&gt;PyPI&lt;&#x2F;a&gt;.
I have been trying to have a deployment that when a git tag is pushed to GitHub,
Travis will test, build and upload both the Anaconda and PyPI packages
for Linux and macOS.&lt;&#x2F;p&gt;
&lt;p&gt;Rather than document what I have found to be the &#x27;perfect&#x27; method,
something which I still don&#x27;t have,
this is a collection of my failures
that will hopefully prevent someone else making the same mistakes.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;packaging-and-deploying&quot;&gt;Packaging and Deploying&lt;&#x2F;h2&gt;
&lt;h3 id=&quot;versioning&quot;&gt;Versioning&lt;&#x2F;h3&gt;
&lt;p&gt;In deploying an application a version scheme provides
an indication of the changes made,
especially where a semantic versioning scheme is used.
To make the versioning process simpler,
I have been using git tags as the &#x27;true&#x27; version number.
To get the version number in the setup
I have been using the &lt;code&gt;setuptools_scm&lt;&#x2F;code&gt; package.
The versioning scheme that best fit how I have been thinking about
release versions is the &lt;a href=&quot;https:&#x2F;&#x2F;www.python.org&#x2F;dev&#x2F;peps&#x2F;pep-0440&#x2F;#post-releases&quot;&gt;post-release&lt;&#x2F;a&gt; scheme,
using&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;use_scm_version={&amp;#x27;version_scheme&amp;#x27;: &amp;#x27;post-release&amp;#x27;}
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;The default state of setuptools_scm is to
only include the dirty flag when files committed in the project have changed,
rather than any untracked file in the directory.
This change to a post release scheme also changed this default,
meaning that when compiling on Travis every commit was dirty.
This was fixed by also specifying the local version scheme&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;use_scm_version={&amp;#x27;version_scheme&amp;#x27;: &amp;#x27;post-release&amp;#x27;,
                 &amp;#x27;local_scheme&amp;#x27;: &amp;#x27;dirty-tag&amp;#x27;},
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;While having a hash after the version number is not a problem for me,
it is a large problem when trying to upload packages to PyPI,
which rejects it with the error&lt;&#x2F;p&gt;
&lt;pre&gt;&lt;code&gt;HTTPError: 400 Client Error: version: Cannot use PEP 440 local versions. for url: https:&amp;#x2F;&amp;#x2F;upload.pypi.org&amp;#x2F;legacy&amp;#x2F;
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;h3 id=&quot;travis&quot;&gt;Travis&lt;&#x2F;h3&gt;
&lt;p&gt;While Travis appears to have good support for uploading to PyPI,
having a &lt;a href=&quot;https:&#x2F;&#x2F;docs.travis-ci.com&#x2F;user&#x2F;deployment&#x2F;pypi&#x2F;&quot;&gt;deployment target&lt;&#x2F;a&gt;,
as far as I can tell it &lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;travis-ci&#x2F;dpl&#x2F;issues&#x2F;377&quot;&gt;doesn&#x27;t support non alphanumeric
passwords&lt;&#x2F;a&gt;.
Unfortunately I was unable to get some alternative methods of including a password to work.
Authentication information for uploading to PyPI can be in the &lt;code&gt;~&#x2F;.pypirc&lt;&#x2F;code&gt; file,
however when trying to write this file using variable expansion&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;bash&quot; class=&quot;language-bash &quot;&gt;&lt;code class=&quot;language-bash&quot; data-lang=&quot;bash&quot;&gt;echo -e &amp;quot;[pypi]\nusername=$PYPI_USERNAME\npassword=$PYPI_PASSWORD&amp;quot; &amp;gt; ~&amp;#x2F;.pypirc
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;resulted in the file&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;conf&quot; class=&quot;language-conf &quot;&gt;&lt;code class=&quot;language-conf&quot; data-lang=&quot;conf&quot;&gt;[pypi]
username=$PYPI_USERNAME
password=$PYPI_PASSWORD
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;with neither of the variables being expanded.
I quickly gave up on that approach
since it worked in every shell I tested it on locally.&lt;&#x2F;p&gt;
&lt;p&gt;A solution that works is to use [twine][twine]
which along with generally being more secure
is able to read the environment variables
&lt;code&gt;TWINE_USERNAME&lt;&#x2F;code&gt; and &lt;code&gt;TWINE_PASSWORD&lt;&#x2F;code&gt; to authenticate.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;installing-from-pip&quot;&gt;Installing from Pip&lt;&#x2F;h3&gt;
&lt;p&gt;Once a package is uploaded to PyPI it needs to be installable,
i.e. running&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$ pip install &amp;lt;package&amp;gt;
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;actually works,
even in a clean environment.
A sample &lt;code&gt;setup.py&lt;&#x2F;code&gt; file for a python package typically looks something like the one below
(&lt;a href=&quot;https:&#x2F;&#x2F;stackoverflow.com&#x2F;questions&#x2F;32528560&#x2F;using-setuptools-to-create-a-cython-package-calling-an-external-c-library&quot;&gt;1&lt;&#x2F;a&gt;,
&lt;a href=&quot;https:&#x2F;&#x2F;stackoverflow.com&#x2F;questions&#x2F;35497572&#x2F;using-python-setuptools-to-put-cython-compiled-pyd-files-in-their-original-folde&quot;&gt;2&lt;&#x2F;a&gt;,
&lt;a href=&quot;http:&#x2F;&#x2F;cython.readthedocs.io&#x2F;en&#x2F;latest&#x2F;src&#x2F;quickstart&#x2F;build.html&quot;&gt;3&lt;&#x2F;a&gt;).&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;# setup.py
from setuptools import setup, Extension
from Cython.Build import cythonize
import numpy

my_extensions = [Extension(..., include_dirs=numpy.get_include())]

setup(
    name=...,
    ...,
    setup_requires=[&amp;#x27;Cython&amp;#x27;, &amp;#x27;numpy&amp;#x27;],
    install_requires=[&amp;#x27;numpy&amp;#x27;]
    ...,
    ext_modules=cythonize(my_extensions),
)
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Problem is, when installing into a new environment via pip
the install fails as reading the &lt;code&gt;setup.py&lt;&#x2F;code&gt; file requires
both Cython and numpy to be installed as they are imported at the top of the file.
This is a significant problem with python packaging 
for which a solution has been accepted &lt;a href=&quot;https:&#x2F;&#x2F;www.python.org&#x2F;dev&#x2F;peps&#x2F;pep-0518&#x2F;&quot;&gt;PEP 518&lt;&#x2F;a&gt; and work is in [progress][pep 518 progress].
In meantime, distributing a wheel is a solution to this problem.&lt;&#x2F;p&gt;
&lt;hr &#x2F;&gt;
&lt;p&gt;&lt;strong&gt;Update 2018-08-12&lt;&#x2F;strong&gt;&lt;&#x2F;p&gt;
&lt;p&gt;The release of pip 10.0 introduced support for a &lt;code&gt;pyproject.toml&lt;&#x2F;code&gt; file in which installation
dependencies can be specified. A &lt;code&gt;pyproject.toml&lt;&#x2F;code&gt; for the above project would look like:&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;toml&quot; class=&quot;language-toml &quot;&gt;&lt;code class=&quot;language-toml&quot; data-lang=&quot;toml&quot;&gt;# pyproject.toml
[build-system]
requires= [&amp;#x27;setuptools&amp;#x27;, &amp;#x27;wheel&amp;#x27;, &amp;#x27;numpy&amp;#x27;, &amp;#x27;Cython&amp;#x27;]
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;hr &#x2F;&gt;
&lt;h2 id=&quot;wheels&quot;&gt;Wheels&lt;&#x2F;h2&gt;
&lt;p&gt;I have included wheels as their own section
since I feel they are especially egregious.
The PyPI &lt;a href=&quot;https:&#x2F;&#x2F;packaging.python.org&#x2F;tutorials&#x2F;distributing-packages&#x2F;#wheels&quot;&gt;documentation&lt;&#x2F;a&gt; tells you that
&amp;quot;You should also create a wheel for your project.&amp;quot;
Well sure, I will just follow the instructions in the docs&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;bash&quot; class=&quot;language-bash &quot;&gt;&lt;code class=&quot;language-bash&quot; data-lang=&quot;bash&quot;&gt;$ python setup.py bdist_wheel
$ twine upload dist&amp;#x2F;sdanalysis-0.4.4.post5-cp36-cp36m-linux_x86_64.whl
Uploading distributions to https:&amp;#x2F;&amp;#x2F;upload.pypi.org&amp;#x2F;legacy&amp;#x2F;
Enter your username: malramsay64
Enter your password:
Uploading sdanalysis-0.4.4.post5-cp36-cp36m-linux_x86_64.whl
HTTPError: 400 Client Error: Binary wheel &amp;#x27;sdanalysis-0.4.4.post5-cp36-cp36m-linux_x86_64.whl&amp;#x27; has an unsupported platform tag &amp;#x27;linux_x86_64&amp;#x27;. for url: https:&amp;#x2F;&amp;#x2F;upload.pypi.org&amp;#x2F;legacy&amp;#x2F;

&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Huh??? 
I also think it is great that I have to wait for the package to upload to get the error message.&lt;&#x2F;p&gt;
&lt;p&gt;Some googling later...&lt;&#x2F;p&gt;
&lt;p&gt;I need to use &lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;pypa&#x2F;auditwheel&quot;&gt;auditwheel&lt;&#x2F;a&gt; to create a &lt;code&gt;manylinux1&lt;&#x2F;code&gt; wheel.
After installing via pip (conda-forge only has auditwheel versions for up to python 3.5)&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;bash&quot; class=&quot;language-bash &quot;&gt;&lt;code class=&quot;language-bash&quot; data-lang=&quot;bash&quot;&gt;$ auditwheel repair dist&amp;#x2F;sdanalysis-0.4.4.post5-cp36-cp36m-linux_x86_64.whl
Repairing sdanalysis-0.4.4.post5-cp36-cp36m-linux_x86_64.whl
usage: auditwheel [-h] [-V] [-v] command ...
auditwheel: error: cannot repair &amp;quot;dist&amp;#x2F;sdanalysis-0.4.4.post5-cp36-cp36m-linux_x86_64.whl&amp;quot; to
&amp;quot;manylinux1_x86_64&amp;quot; ABI because of the presence of too-recent versioned symbols. You&amp;#x27;ll need to
compile the wheel on an older toolchain.

&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Some more googling and reading documentation ...&lt;&#x2F;p&gt;
&lt;p&gt;From the auditwheel README:&lt;&#x2F;p&gt;
&lt;blockquote&gt;
&lt;p&gt;But in general, building manylinux1 wheels requires running on a CentOS5 machine, so we recommend using the pre-built manylinux Docker image.&lt;&#x2F;p&gt;
&lt;&#x2F;blockquote&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$ docker run -i -t -v `pwd`:&amp;#x2F;io quay.io&amp;#x2F;pypa&amp;#x2F;manylinux1_x86_64 &amp;#x2F;bin&amp;#x2F;bash
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Oh easy, now I have to set this up to build in a docker container...&lt;&#x2F;p&gt;
&lt;p&gt;For a great example of building a manylinux1 wheel,
I can suggest checking out the &lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;pypa&#x2F;python-manylinux-demo&quot;&gt;pypa&#x2F;python-manylinux-demo&lt;&#x2F;a&gt; repository.
Unfortunately it is not linked from the python packaging documentation
and so is not that simple to find.&lt;&#x2F;p&gt;
&lt;p&gt;Packaging in python is hard.
I have over 70 builds this week on travis
just trying to get everything to actually work.
At the end of all this I still don&#x27;t really know what I am doing
other than basically everything I try doesn&#x27;t work
and that I shouldn&#x27;t try any of the above,
since I already know it won&#x27;t work.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;building-and-testing&quot;&gt;Building and Testing&lt;&#x2F;h2&gt;
&lt;p&gt;A crucial aspect of deploying a piece of software is
having some confidence that it actually works,
even if that just means it doesn&#x27;t break on import.
For my workflow this is an automated process,
with builds running on &lt;a href=&quot;https:&#x2F;&#x2F;travis-ci.org&#x2F;&quot;&gt;Travis-CI&lt;&#x2F;a&gt; when
a new commit is pushed to Github.
Even using a popular service live Travis didn&#x27;t preclude me
from having plenty of failures.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;travis-python-and-macos&quot;&gt;Travis, Python and macOS&lt;&#x2F;h3&gt;
&lt;p&gt;I had automated testing set up on Travis-CI for my python module on linux,
however trying to extent this to also using macOS 
I found some unusual behaviour.
At first I thought that testing on travis was going to be
as simple as adding it to the list of os&#x27;s.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;yaml&quot; class=&quot;language-yaml &quot;&gt;&lt;code class=&quot;language-yaml&quot; data-lang=&quot;yaml&quot;&gt;# .travis.yml
language: python

os:
  - linux
  - osx
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;However it really wasn&#x27;t because the build raised an error
before my code even started running.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$ sudo tar xjf python-3.6.tar.bz2 --directory &amp;#x2F;
tar: Unrecognized archive format
tar: Error exit delayed from previous errors.
The command &amp;quot;sudo tar xjf python-3.6.tar.bz2 --directory &amp;#x2F;&amp;quot; failed and exited with 1 during .
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;It turns out that on Travis-CI,
the python language is &lt;a href=&quot;https:&#x2F;&#x2F;docs.travis-ci.com&#x2F;user&#x2F;multi-os&#x2F;#Python-example-(unsupported-languages)&quot;&gt;unsupported on macOS&lt;&#x2F;a&gt;.
So I had to bodge some kind of workaround,
especially since the &lt;a href=&quot;https:&#x2F;&#x2F;docs.travis-ci.com&#x2F;user&#x2F;multi-os&#x2F;#Python-example-(unsupported-languages)&quot;&gt;travis docs&lt;&#x2F;a&gt; on the issue
assume you are using &lt;a href=&quot;https:&#x2F;&#x2F;tox.readthedocs.io&#x2F;en&#x2F;latest&#x2F;&quot;&gt;tox&lt;&#x2F;a&gt;.
I found were that when using conda to install python,
the simplest workaround is to use &lt;code&gt;language: generic&lt;&#x2F;code&gt;
and install python through conda for each OS.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;yaml&quot; class=&quot;language-yaml &quot;&gt;&lt;code class=&quot;language-yaml&quot; data-lang=&quot;yaml&quot;&gt;# .travis.yml
language: generic

os:
  - linux
  - osx

before_install:
  - if [[ $(uname -s) == &amp;quot;Darwin&amp;quot; ]]; then
      wget https:&amp;#x2F;&amp;#x2F;repo.continuum.io&amp;#x2F;miniconda&amp;#x2F;Miniconda3-latest-MacOSX-x86_64.sh -O miniconda.sh;
    else
      wget https:&amp;#x2F;&amp;#x2F;repo.continuum.io&amp;#x2F;miniconda&amp;#x2F;Miniconda3-latest-Linux-x86_64.sh -O miniconda.sh;
    fi
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Alternatively, when just using the latest version of python
it can be installed using homebrew.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;yaml&quot; class=&quot;language-yaml &quot;&gt;&lt;code class=&quot;language-yaml&quot; data-lang=&quot;yaml&quot;&gt;# .travis.yml
language: python

matrix:
  include:
    - os: linux
      python: 3.6

    - os: osx
      language: generic
      before_install: brew install python3
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;h3 id=&quot;frozen-versions-melt&quot;&gt;Frozen versions melt&lt;&#x2F;h3&gt;
&lt;p&gt;When deploying with [Conda][conda] the recipe is created in a [meta.yaml][meta.yaml] file,
containing the requirements for each step,
in addition to instructions for building and testing of a package.
A very simple version is shown below specifying a patch version of python (&lt;code&gt;3.6.1&lt;&#x2F;code&gt;) 
to use for both the build and the run phase.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;yaml&quot; class=&quot;language-yaml &quot;&gt;&lt;code class=&quot;language-yaml&quot; data-lang=&quot;yaml&quot;&gt;# file: meta.yaml
package:
  name: pyver_test
  version: 1.0.0

requirements:
  build:
    - python 3.6.1
  run:
    - python {{ pin_compatible(&amp;#x27;python&amp;#x27;, max_pin=&amp;#x27;x.x.x&amp;#x27;) }}

build: number: 0
test:
  commands: python --version
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;While the version number is adhered to in the build phase,
during the run phase the patch version is ignored,
instead installing the latest patch version of python,
which at the time of writing is 3.6.4.
This is the &lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;conda&#x2F;conda-build&#x2F;issues&#x2F;2571&quot;&gt;intended behaviour&lt;&#x2F;a&gt;,
albeit somewhat unexpected.&lt;&#x2F;p&gt;
&lt;p&gt;The reason that this was a problem for me is that
I came across an &lt;a href=&quot;https:&#x2F;&#x2F;bitbucket.org&#x2F;glotzer&#x2F;freud&#x2F;issues&#x2F;154&#x2F;conda-package-broken-on-macos-with-python#comment-41977740&quot;&gt;unusual bug&lt;&#x2F;a&gt; in a conda package
I was trying to use.
When installed on macOS&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$  conda create -n freud -c glotzer freud
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;importing the module would crash python&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$ python -c &amp;quot;import freud&amp;quot;
Fatal Python error: PyThreadState_Get: no current thread

zsh: abort           python -c &amp;quot;import freud&amp;quot;
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;This environment would use the latest patch version of python (currently 3.6.4)
and a simple workaround was to install &lt;code&gt;python==3.6.1&lt;&#x2F;code&gt;,
although it did make installing the package more difficult than I would have liked.
I don&#x27;t know why this worked, it just did,
although I suspect it has something to do with linking to the python C libraries.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;when-the-porcelain-stains&quot;&gt;When the porcelain stains&lt;&#x2F;h3&gt;
&lt;p&gt;&lt;a href=&quot;https:&#x2F;&#x2F;pypi.python.org&#x2F;pypi&#x2F;pipenv&quot;&gt;Pipenv&lt;&#x2F;a&gt; is a packaging manager for python that automates
the creation of virtual environments for python projects
in addition to the packages installed within them.
As Jannis Leidel is quoted in the testimonials;
&amp;gt;Pipenv is the porcelain I always wanted to build for pip.
However there are always unusual cases when abstracting that can cause issues.&lt;&#x2F;p&gt;
&lt;p&gt;The first of these issues I found was in the installation,
I primarily use Conda for my python and virtual environment configuration,&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$ pip install pipenv
$ pipenv install numpy
ERROR: virtualenv is not compatible with this system or executable
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;However, it is required to install pipenv outside of a conda distribution
(&lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;pypa&#x2F;pipenv&#x2F;issues&#x2F;699&quot;&gt;issue#699&lt;&#x2F;a&gt;, &lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;pypa&#x2F;pipenv&#x2F;issues&#x2F;288&quot;&gt;issue#288&lt;&#x2F;a&gt;).
So installation can be performed by running&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$ &amp;#x2F;usr&amp;#x2F;local&amp;#x2F;bin&amp;#x2F;pip install pipenv
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;or with conda 4.4&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$ conda deactivate
$ pip install pipenv
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;One of the major features of Pipenv over other package managers for python
is that it includes hashes of the binaries as part of the version pinning.
This both ensures you are installing exactly the same code on different systems,
and provides security in preventing unknown changes from being installed.
These hashes are collected for all the package versions available on PyPI,
allowing for cross platform and python version use.&lt;&#x2F;p&gt;
&lt;p&gt;The problem that I came across is that a certain package &lt;a href=&quot;https:&#x2F;&#x2F;pypi.python.org&#x2F;pypi&#x2F;numpy-quaternion&quot;&gt;numpy-quaternion&lt;&#x2F;a&gt;
has some nearly identical version numbers,&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;numpy-quaternion 2017.11.26.19.0.13&lt;&#x2F;code&gt;&lt;&#x2F;li&gt;
&lt;li&gt;&lt;code&gt;numpy-quaternion 2017.11.26.19.00.13&lt;&#x2F;code&gt;&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;the second just having an additional zero.
The first of these version numbers contains all the pre-built wheels,
while the second contains the source.
When downloading the hashes,
pipenv downloads the hash for the source distribution
rather than for the wheels.
This means that when trying to install in a new virtualenv
the downloaded package fails the hash comparison&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;txt&quot; class=&quot;language-txt &quot;&gt;&lt;code class=&quot;language-txt&quot; data-lang=&quot;txt&quot;&gt;THESE PACKAGES DO NOT MATCH THE HASHES FROM Pipfile.lock!
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;and aborts the install,
incredibly frustrating when the initial install of the package works fine.&lt;&#x2F;p&gt;
&lt;p&gt;A workaround is to manually update the Pipfile.lock file
with the correct hashes,
which is fine while dependencies are not updated 
and the Pipfile.lock file remains unchanged.&lt;&#x2F;p&gt;
&lt;hr &#x2F;&gt;
&lt;p&gt;Update 2018-01-09&lt;&#x2F;p&gt;
&lt;p&gt;In a previous revision of the Installing from Pip section
I stated that importing the dependencies within a function call
was a solution to the import problems.
Thanks to &lt;a href=&quot;https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;Python&#x2F;comments&#x2F;7ope3j&#x2F;a_guide_on_how_not_to_package_a_python_module&#x2F;dscmwf5&#x2F;&quot;&gt;&#x2F;u&#x2F;vorpalsmith&lt;&#x2F;a&gt;
for bringing it to my attention that this isn&#x27;t a solution,
and that the success that I observed was a result of installing a wheel
rather than any changes to the code that I made.&lt;&#x2F;p&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>Quaternions for Orientation in MD</title>
        <published>2017-09-13T00:00:00+00:00</published>
        <updated>2017-09-13T00:00:00+00:00</updated>
        <author>
          <name>Unknown</name>
        </author>
        <link rel="alternate" href="https://malramsay.com/post/quaternions-in-md/" type="text/html"/>
        <id>https://malramsay.com/post/quaternions-in-md/</id>
        
        <content type="html">&lt;p&gt;In the understanding of the dynamics of a molecule in a Molecular Dynamics (MD) Simulation
the two main properties we use to understand liquid behaviour are 
the translational motion and the rotational motion of the atoms or molecules.
These motions reflect changes in the position and the orientation of the molecules.
When calculating these motions in an MD simulation
it makes sense to represent each molecule as;&lt;&#x2F;p&gt;
&lt;ul&gt;
&lt;li&gt;the position of the Center of Mass (COM), and&lt;&#x2F;li&gt;
&lt;li&gt;the orientation within the lab reference frame.&lt;&#x2F;li&gt;
&lt;&#x2F;ul&gt;
&lt;p&gt;This representation allows for quick and simple calculation of 
the motion between two separate times.
In the simplest case, conceptually this is just computing&lt;&#x2F;p&gt;
&lt;p&gt;\[
\begin{aligned}
\text{displacement} &amp;amp;= \text{norm}(\text{position}_2 - \text{position}_1) \\
\text{rotation} &amp;amp;= \text{norm}(\text{orientation}_2 - \text{orientation}_1).
\end{aligned}
\]&lt;&#x2F;p&gt;
&lt;p&gt;For the displacement,
we can just use a standard Euclidean distance norm.
Which is simple enough to compute.
The rotations are somewhat more difficult.
In a two dimensional system it would be possible to use just the angle,
where the norm is a function keeping the rotation within a sensible bounds, like $(-\pi, \pi]$.
However in three dimensions using angles gets more complicated
and computationally expensive.
The approach that is used in &lt;a href=&quot;http:&#x2F;&#x2F;glotzerlab.engin.umich.edu&#x2F;hoomd-blue&#x2F;&quot;&gt;hoomd&lt;&#x2F;a&gt; 
and likely many other molecular dynamics programs is 
to use quaternions for the representation of orientation.&lt;&#x2F;p&gt;
&lt;p&gt;This post aims to give an understanding of quaternions
and how to use them to compute a rotation
from one orientation to another.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;orientation-in-2d&quot;&gt;Orientation in 2D&lt;&#x2F;h2&gt;
&lt;p&gt;We will start with a simple case,
orientation in two dimensions.
In two dimensions it is possible to describe the orientation of a molecules
using a single angle, $\theta$.
This is a perfectly reasonable representation.
We can find the rotational distance $\Delta\theta$ between two orientations as&lt;&#x2F;p&gt;
&lt;p&gt;\[
\Delta\theta = \theta_2 - \theta_1
\]&lt;&#x2F;p&gt;
&lt;p&gt;This can become a bit of an issue if we want to keep our angle bounded.
A reasonable range for an angle of rotation is $(-\pi,pi]$,
as we can assume a rotation larger occurred in the opposite direction.
The code to implement this in python would look something like the function below;&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;import math

def rotationalDistance(theta1, theta2):
    delta_theta = theta2 - theta1
    if delta_theta &amp;gt; math.pi:
        delta_theta -= 2*math.pi
    elif delta_theta &amp;lt;= -math.pi:
        delta_theta += 2*math.pi
    return delta_theta
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;h3 id=&quot;complex-numbers&quot;&gt;Complex Numbers&lt;&#x2F;h3&gt;
&lt;p&gt;An equally valid method for the representation of angle in 2D is to use complex numbers.
Any complex number $z = a+ib$ can be represented in exponential form&lt;&#x2F;p&gt;
&lt;p&gt;\[
z = r\text{e}^{i\theta}
\]&lt;&#x2F;p&gt;
&lt;p&gt;where $r$ is the length, or modulo, of $z$&lt;&#x2F;p&gt;
&lt;p&gt;\[
r = \sqrt{a^2 + b^2}
\]&lt;&#x2F;p&gt;
&lt;p&gt;and $\theta$ is the argument of $z$,
which for the first quadrant is&lt;&#x2F;p&gt;
&lt;p&gt;\[
\theta = \tan^{-1}\frac{b}{a}
\]&lt;&#x2F;p&gt;
&lt;p&gt;Due to the range of the $\tan^{-1}$ function only being $[-\pi&#x2F;2, \pi&#x2F;2]$,
the quadrant of the complex number is important to having values in the range $(-\pi,pi]$.
There is a nice computational solution to this problem implemented in almost all languages,
the &lt;code&gt;atan2&lt;&#x2F;code&gt; function which works out the quadrant for us,
giving an angle in our desired range.&lt;&#x2F;p&gt;
&lt;p&gt;\[
\theta = \text{atan2}(b, a)
\]&lt;&#x2F;p&gt;
&lt;p&gt;Make special note of the order of arguments to the &lt;code&gt;atan2&lt;&#x2F;code&gt; function.&lt;&#x2F;p&gt;
&lt;p&gt;Now that we understand how to convert our final complex number to a rotation,
we need to construct it from the distance between two orientations.
When we multiply two complex numbers together it can be thought of as rotating by an angle.
This can be demonstrated by performing a multiplication with 
the polar form of a complex number.&lt;&#x2F;p&gt;
&lt;p&gt;\[
\begin{aligned}
z &amp;amp;= r_1e^{i\theta_1} \times r_2e^{i\theta_2}\
&amp;amp;  = r_1r_2e^{i\theta_1 + i\theta_2}\
&amp;amp;  = r_1r_2e^{i(\theta_1 + \theta_2)}
\end{aligned}
\]&lt;&#x2F;p&gt;
&lt;p&gt;Conversely division of a complex number by another is the distance between the two angles&lt;&#x2F;p&gt;
&lt;p&gt;\[
\begin{aligned}
z &amp;amp;= r_1e^{i\theta_1} &#x2F; r_2e^{i\theta_2} \
&amp;amp;= \frac{r_1}{r_2}e^{i\theta_1 - i\theta_2} \
&amp;amp;= \frac{r_1}{r_2}e^{i(\theta_1 - \theta_2)}
\end{aligned}
\]&lt;&#x2F;p&gt;
&lt;p&gt;It is possible to do all our intermediate calculations using 
the mathematics of complex numbers.
Then when we need an rotational distance
the complex number representing that can be converted to a angle.&lt;&#x2F;p&gt;
&lt;p&gt;Implementing this is produces a function that looks like the one below,&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;import math

def complexRotation(initial, final):
    delta = final&amp;#x2F;initial
    return math.atan2(delta.imag, delta.real)
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;While the code is written in python,
most programming languages will handle complex numbers in a similar fashion.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;orientation-in-3d&quot;&gt;Orientation in 3D&lt;&#x2F;h2&gt;
&lt;p&gt;Dealing with the 2D case was relatively simple.
There was a single orientation to deal with,
you are probably somewhat familiar with complex numbers
(or at least were at some point),
and it is simple to draw and visualise.
In 3D things are a little more difficult.
When expressing orientation using rotations in 3D,
known as Euler angles,
there is a phenomenon known as gimbal lock.
This is where you lose a degree of freedom of expressing orientation,
similar to in spherical coordinates where if you have $\theta = 0$
then all values of $\phi$ are equivalent i.e. that degree of freedom is lost.
This means it is typically a bad idea to represent 3D orientation using angles.
Instead we use the equivalent of complex numbers, quaternions.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;quaternions&quot;&gt;Quaternions&lt;&#x2F;h3&gt;
&lt;p&gt;Just like complex numbers,
quaternions are a mathematical construct that allow us to represent rotations,
only in three dimensions instead of two.
As the name suggests quaternions consist of four values of the form&lt;&#x2F;p&gt;
&lt;p&gt;\[
a + b\mathbf{i} + c\mathbf{j} + d\mathbf{k}
\]&lt;&#x2F;p&gt;
&lt;p&gt;where the relationships&lt;&#x2F;p&gt;
&lt;p&gt;\[
\mathbf{i}^2 = \mathbf{j}^2 = \mathbf{k}^2 = \mathbf{ijk} = -1
\]&lt;&#x2F;p&gt;
&lt;p&gt;hold true.
These relationships are essentially giving us the rules for
the operations on the quaternions.&lt;&#x2F;p&gt;
&lt;p&gt;These operations on quaternions are very similar to those on complex numbers.
The conjugate of a quaternion $q$ is;&lt;&#x2F;p&gt;
&lt;p&gt;\[
q^* = a - b\mathbf{i} - c\mathbf{j} - d\mathbf{k}
\]&lt;&#x2F;p&gt;
&lt;p&gt;The norm of $q$ is&lt;&#x2F;p&gt;
&lt;p&gt;\[
||q|| = \sqrt{q^*q} = \sqrt{qq^*} = \sqrt{a^2 + b^2 + c^2 + d^2}
\]&lt;&#x2F;p&gt;
&lt;p&gt;The reciprocal of $q$ is given as&lt;&#x2F;p&gt;
&lt;p&gt;\[
q^{-1} = \frac{q^*}{||a||^2}
\]&lt;&#x2F;p&gt;
&lt;p&gt;Quaternions have the same properties of multiplication and division as complex numbers.
Multiplying a quaternion by another is applying a rotation,
and the division is finding the rotation from one quaternion to the other.&lt;&#x2F;p&gt;
&lt;h3 id=&quot;rotational-magnitude&quot;&gt;Rotational Magnitude&lt;&#x2F;h3&gt;
&lt;p&gt;There are a number of methods for computing rotational magnitude using quaternions,
some of which trade of speed of calculation for accuracy.
For an overview of the different methods of computing rotational distance Du Q. Huynh&lt;sup class=&quot;footnote-reference&quot;&gt;&lt;a href=&quot;#1&quot;&gt;1&lt;&#x2F;a&gt;&lt;&#x2F;sup&gt;
does an excellent comparison of six different methods.&lt;&#x2F;p&gt;
&lt;p&gt;Of the methods used to compute rotational magnitude $\Phi$,
the one I think is most suitable for use in Molecular Dynamics is&lt;&#x2F;p&gt;
&lt;p&gt;[
\Phi = 2\ \text{arccos}(|\mathbf{q}_1 · \mathbf{q}_2|)
]&lt;&#x2F;p&gt;
&lt;p&gt;This gives a value on the range $[0, 2\pi]$ in units of radians.
I think this is most suitable as it gives results in units of radians,
rather than some approximate distance,
while remaining simple and fast to compute.
Implemented in python this function is;&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;import numpy

def quaternion_distance(initial, final):
    return 2*numpy.arccos(numpy.abs(numpy.dot(initial, final)))
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;For an optimised computation in python over an array of values,
the below is the fastest implementation I could find.&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;import numpy
def quaternion_distance_array(initial, final):
    return 2*numpy.arccos(numpy.abs(numpy.einsum(&amp;#x27;ij,ij-&amp;gt;i&amp;#x27;, initial, final)))
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;For more complicated operations using quaternions I would suggest having a look at &lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;moble&#x2F;quaternion&quot;&gt;quaternion&lt;&#x2F;a&gt;,
an open source python module by Mike Boyle which adds support for quaternions to numpy.
Mike also has a &lt;a href=&quot;https:&#x2F;&#x2F;github.com&#x2F;moble&#x2F;Quaternions&quot;&gt;C++&lt;&#x2F;a&gt; version of the library available.&lt;&#x2F;p&gt;
&lt;div class=&quot;footnote-definition&quot; id=&quot;1&quot;&gt;&lt;sup class=&quot;footnote-definition-label&quot;&gt;1&lt;&#x2F;sup&gt;
&lt;p&gt;Huynh, D. Q. (2009). Metrics for 3D rotations: Comparison and analysis. Journal of Mathematical Imaging and Vision, 35(2), 155–164. &lt;a href=&quot;https:&#x2F;&#x2F;doi.org&#x2F;10.1007&#x2F;s10851-009-0161-2&quot;&gt;doi: 0.1007&#x2F;s10851-009-0161-2&lt;&#x2F;a&gt; (&lt;a href=&quot;http:&#x2F;&#x2F;ai2-s2-pdfs.s3.amazonaws.com&#x2F;5617&#x2F;8de1001efe54792ad93f6980de5d5e91906b.pdf&quot;&gt;#icanhazpdf&lt;&#x2F;a&gt;)&lt;&#x2F;p&gt;
&lt;&#x2F;div&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>Building interfaces with click</title>
        <published>2017-08-03T00:00:00+00:00</published>
        <updated>2017-08-03T00:00:00+00:00</updated>
        <author>
          <name>Unknown</name>
        </author>
        <link rel="alternate" href="https://malramsay.com/post/building-with-click/" type="text/html"/>
        <id>https://malramsay.com/post/building-with-click/</id>
        
        <content type="html">&lt;p&gt;Running computer simulations is great.
You tell the computer what to do,
it sits there crunching numbers for a while.
Then you can come back and look at the results.
The problem with this is that most simulation packages
take the simulation parameters in an input file,
so running a simulation becomes;&lt;&#x2F;p&gt;
&lt;ol&gt;
&lt;li&gt;Write an input file&lt;&#x2F;li&gt;
&lt;li&gt;Run simulation&lt;&#x2F;li&gt;
&lt;li&gt;Edit input file&lt;&#x2F;li&gt;
&lt;li&gt;Run simulation&lt;&#x2F;li&gt;
&lt;li&gt;Forget what the old file looked like&lt;&#x2F;li&gt;
&lt;&#x2F;ol&gt;
&lt;p&gt;Click can help us break this pattern,
providing us with a simple way to pass values to scripts from the command line.
Allowing us to build command line applications that 
take care of editing these input files for us.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;what-is-click&quot;&gt;What is Click?&lt;&#x2F;h2&gt;
&lt;p&gt;Click is a python library that provides function decorators for
passing arguments to function calls through arguments and options on the command line.
Click handles all the parsing for us, including the parsing of types,
and also allows for some simple validation of values.&lt;&#x2F;p&gt;
&lt;p&gt;Understanding click is probably best done through an example.
Let&#x27;s say we have a script to run a simulation that looks like below&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;#!&amp;#x2F;usr&amp;#x2F;bin&amp;#x2F;env python
# run_simulation.py

import time

def main():
    time.sleep(30)
    print(&amp;#x27;Simulation Finished!&amp;#x27;)

if __name__ == &amp;#x27;__main__&amp;#x27;:
    main()
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;To change the length of time we seep for we could edit the script manually,
but we want to sleep for a whole range of different times.
To pass an sleep value from the command line we can modify the script as follows;&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;#!&amp;#x2F;usr&amp;#x2F;bin&amp;#x2F;env python
# run_simulation.py

import time
import click

@click.command()
@click.argument(&amp;#x27;sleep_time&amp;#x27;, type=int)
def main(sleep_time):
    time.sleep(sleep_time)
    print(&amp;#x27;Simulation Finished!&amp;#x27;)

if __name__ == &amp;#x27;__main__&amp;#x27;:
    main()
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;The script can then be run as follows&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$ python run_simulation.py 1
Simulation Finished!
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;which will sleep for 1 second then print the output.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;taking-it-up-a-level&quot;&gt;Taking it up a level&lt;&#x2F;h2&gt;
&lt;p&gt;Changing the time we sleep for is nice,
but we are lazy and don&#x27;t want to have to specify this every time.
We also want to be able to specify the output string 
and the file the output is written to.
Easy!!!&lt;&#x2F;p&gt;
&lt;p&gt;To make the &lt;code&gt;sleep_time&lt;&#x2F;code&gt; optional we can use the following like of code &lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;@click.option(&amp;#x27;-t&amp;#x27;, &amp;#x27;--sleep-time&amp;#x27;, type=int, default=5)
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;now rather than passing the time as an ordered argument,
we pass it as an option as either the short &lt;code&gt;-t &amp;lt;time&amp;gt;&lt;&#x2F;code&gt; 
or long &lt;code&gt;--sleep-time &amp;lt;time&amp;gt;&lt;&#x2F;code&gt; way.
Note that the variable name is extracted from the long form
with the hyphen being converted to an underscore.&lt;&#x2F;p&gt;
&lt;p&gt;There is another improvement we can make to this &lt;code&gt;sleep_time&lt;&#x2F;code&gt; option.
It is impossible to sleep for a negative time,
and we don&#x27;t want to wait for longer than a minute for this to run.
Click allows us to set a range of suitable values&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;@click.option(&amp;#x27;-t&amp;#x27;, &amp;#x27;--sleep-time&amp;#x27;, type=click.IntRange(min=0, max=60), default=5)
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Now when we run the simulation&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$ python run_simulation.py
Simulation Finished!
$ python run_simulation.py -t 5
Simulation Finished!
$ python run_simulation.py -t -1
Usage: run_simulation.py [OPTIONS]

Error: Invalid value for &amp;quot;-t&amp;quot; &amp;#x2F; &amp;quot;--sleep-time&amp;quot;: -1 is not in the valid range of 0 to 60.
$ python run_simulation.py -t 61
Usage: run_simulation.py [OPTIONS]

Error: Invalid value for &amp;quot;-t&amp;quot; &amp;#x2F; &amp;quot;--sleep-time&amp;quot;: 61 is not in the valid range of 0 to 60.
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Awesome, we can stop ourselves accidentally running a simulation that takes forever
and we get a helpful error message when we do.
In fact click will generate help documentation for us&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$ python run_simulation.py --help
Usage: run_simulation.py [OPTIONS]

Options:
  -t, --sleep-time INTEGER RANGE
  --help                          Show this message and exit.
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;We can help it generate this help by providing a help string to our arguments and options&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;@click.option(&amp;#x27;-t&amp;#x27;, &amp;#x27;--sleep-time&amp;#x27;, type=click.IntRange(min=0, max=60), default=5,
              help=&amp;#x27;Specify the time for which the simulation will run.&amp;#x27;)
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;So we have the sleep time sorted,
now for the output string and file.
The output string is just like the sleep time,
we can add another option;&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;@click.option(&amp;#x27;--outstring&amp;#x27;, type=str, default=&amp;#x27;Simulation Finished!&amp;#x27;,
              help=&amp;#x27;String to print to output file at the end of the simulation.&amp;#x27;)
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;Just make sure that when you pass a value to this option 
that spaces are either escaped or quoted, 
otherwise they will be parsed as separate arguments.&lt;&#x2F;p&gt;
&lt;p&gt;Click also has excellent support for file input and output,
including reading from stdin and writing to stdout.
We get this functionality by using  &lt;code&gt;click.File&lt;&#x2F;code&gt; as the type like below&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;@click.argument(&amp;#x27;output&amp;#x27;, default=&amp;#x27;-&amp;#x27;, type=click.File(&amp;#x27;w&amp;#x27;))
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;What does this all mean.
There is an argument, &lt;code&gt;output&lt;&#x2F;code&gt; which has the type of a writable file.
It will return a file object that we can write to directly,
click handles the open&#x2F;close for us.
The default value &lt;code&gt;&#x27;-&#x27;&lt;&#x2F;code&gt; is stdout,
like in other command line tools where you can use &lt;code&gt;-&lt;&#x2F;code&gt; to indicate
read from or write to stdin&#x2F;stdout,
click has the same support.&lt;&#x2F;p&gt;
&lt;h2 id=&quot;bringing-it-all-together&quot;&gt;Bringing it all together&lt;&#x2F;h2&gt;
&lt;p&gt;With all these modifications we end up with a script that looks like this;&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;python&quot; class=&quot;language-python &quot;&gt;&lt;code class=&quot;language-python&quot; data-lang=&quot;python&quot;&gt;#!&amp;#x2F;usr&amp;#x2F;bin&amp;#x2F;env python
# run_simulation.py

import time
import click

@click.command()
@click.argument(&amp;#x27;output&amp;#x27;, default=&amp;#x27;-&amp;#x27;, type=click.File(&amp;#x27;w&amp;#x27;))
@click.option(&amp;#x27;-t&amp;#x27;, &amp;#x27;--sleep-time&amp;#x27;, type=click.IntRange(min=0, max=60), default=5,
              help=&amp;#x27;Specify the time for which the simulation will run.&amp;#x27;)
@click.option(&amp;#x27;--outstring&amp;#x27;, type=str, default=&amp;#x27;Simulation Finished!&amp;#x27;,
              help=&amp;#x27;String to print to OUTPUT at the end of the simulation.&amp;#x27;)
def main(output, sleep_time, outstring):
    time.sleep(sleep_time)
    output.write(outstring + &amp;#x27;\n&amp;#x27;)

if __name__ == &amp;#x27;__main__&amp;#x27;:
    main()
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;To work out how to run the script we can use the help&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$ python run_simulation.py --help
Usage: run_simulation.py [OPTIONS] [OUTPUT]

Options:
 -t, --sleep-time INTEGER RANGE  Specify the time for which the simulation
                                 will run.
 --outstring TEXT                String to print to output file at the end of
                                 the simulation.
 --help                          Show this message and exit.
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;or just run the command&lt;&#x2F;p&gt;
&lt;pre data-lang=&quot;sh&quot; class=&quot;language-sh &quot;&gt;&lt;code class=&quot;language-sh&quot; data-lang=&quot;sh&quot;&gt;$ python run_simulation.py -t 1 --outstring &amp;quot;Hello World.&amp;quot; hello.txt
&lt;&#x2F;code&gt;&lt;&#x2F;pre&gt;
&lt;p&gt;We have taken a simple script that can perform a single task
and with an extra 7 lines of code have turned it into a versatile tool.
There is plenty more functionality available with Click,
and you should check out the &lt;a href=&quot;http:&#x2F;&#x2F;click.pocoo.org&#x2F;5&#x2F;&quot;&gt;documentation&lt;&#x2F;a&gt;.&lt;&#x2F;p&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>Shear melting at the crystal-liquid interface: Erosion and the asymmetric suppression of interface fluctuations</title>
        <published>2016-02-16T00:00:00+00:00</published>
        <updated>2016-02-16T00:00:00+00:00</updated>
        <author>
          <name>Ramsay, Malcolm</name>
        </author>
        <link rel="alternate" href="https://malramsay.com/publication/shear-melting/" type="text/html"/>
        <id>https://malramsay.com/publication/shear-melting/</id>
        
        <content type="html">&lt;p&gt;The influence of an applied shear on the planar crystal-melt interface is modeled by a nonlinear stochastic partial differential equation of the interface fluctuations. A feature of this theory is the asymmetric destruction of interface fluctuations due to advection of the crystal protrusions on the liquid side of the interface only. We show that this model is able to qualitatively reproduce the nonequilibrium coexistence line found in simulations. The impact of shear on spherical clusters is also addressed.&lt;&#x2F;p&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>Packing concave molecules in crystals and amorphous solids: On the connection between shape and local structure</title>
        <published>2015-04-21T00:00:00+00:00</published>
        <updated>2015-04-21T00:00:00+00:00</updated>
        <author>
          <name>Jennings, Cerridwen</name>
        </author>
        <link rel="alternate" href="https://malramsay.com/publication/packing-molecules/" type="text/html"/>
        <id>https://malramsay.com/publication/packing-molecules/</id>
        
        <content type="html">&lt;p&gt;The structure of the densest crystal packings is determined for a variety of concave shapes in 2D constructed by the overlap of two or three discs. The maximum contact number per particle pair is defined and proposed as a useful means of categorizing particle shape. We demonstrate that the densest packed crystal exhibits a maximum in the number of contacts per particle but does not necessarily include particle pairs with the maximum contact number. In contrast, amorphous structures, generated by energy minimisation of high temperature liquids, typically do include maximum contact pairs. The amorphous structures exhibit a large number of contacts per particle corresponding to over-constrained structures. Possible consequences of this over-constraint are discussed.&lt;&#x2F;p&gt;
</content>
        
    </entry>
    <entry xml:lang="en">
        <title>Defect-mediated relaxation in the random tiling phase of a binary mixture: Birth, death and mobility of an atomic zipper</title>
        <published>2014-02-14T00:00:00+00:00</published>
        <updated>2014-02-14T00:00:00+00:00</updated>
        <author>
          <name>Tondl, Elizabeth</name>
        </author>
        <link rel="alternate" href="https://malramsay.com/publication/defect-relaxation/" type="text/html"/>
        <id>https://malramsay.com/publication/defect-relaxation/</id>
        
        <content type="html">&lt;p&gt;This paper describes the mechanism of defect-mediated relaxation in a dodecagonal square-triangle random tiling phase exhibited by a simulated binary mixture of soft discs in 2D. We examine the internal transitions within the elementary mobile defect (christened the &#x27;zipper&#x27;) that allow it to move, as well as the mechanisms by which the zipper is created and annihilated. The structural relaxation of the random tiling phase is quantified and we show that this relaxation is well described by a model based on the distribution of waiting times for each atom to be visited by the diffusing zipper. This system, representing one of the few instances where a well defined mobile defect is capable of structural relaxation, can provide a valuable test case for general theories of relaxation in complex and disordered materials.&lt;&#x2F;p&gt;
</content>
        
    </entry>
</feed>
