Rowan Das Résumé
← Back to the main page

Deals aggregator

Who is buying whom

The M&A headlines of the last few days, read each weekday by a Llama model and grouped by deal. Each deal shows where it stands, and the small numbers link to the headline behind every claim.

Deals in play

The best covered deals first. A deal shows once, however many headlines reported it.

Sources

Every headline the model was given, numbered the way it cites them.

    How this page was made

    This is the same retrieval-augmented generation (RAG) design as the morning brief, pointed at one question: who is buying whom? A language model on its own knows nothing about this week's deals, and if you ask it anyway it will make one up. RAG searches for the relevant news first, then hands only that news to the model with instructions to write from it and cite it.

    The brief knows in advance what to look for: six instruments. Deals are different, because nobody tells the pipeline which companies are about to merge. So there is an extra step, between storing the headlines and searching them, where the model reads each headline and works out which deal it is about. Code does everything that can be checked, and the model does the reading and writing. Both run through Ollama on the same machine as the code, so no text goes to an outside AI service.

    1. Step 1

      Pull deal headlines

      Python reads Google News RSS feeds for twelve searches worded the way deal stories are worded: "agrees to acquire", "in talks to acquire", "antitrust review", "merger terminated" and so on. Between them they catch rumors, signed deals, regulatory reviews, closings and collapses. The code parses each item, drops anything older than 72 hours, strips the HTML out of the snippet and removes stories that came back under more than one search.

      Deals move more slowly than markets, so the window is three days, not one. Each story gets an ID made from a hash of its link, so the same article is never stored twice.

      # headlines.py "id": hashlib.sha1(link.encode()).hexdigest()[:16]

    2. Step 2

      Turn each headline into a vector

      An embedding model, nomic-embed-text, reads each headline and its snippet and returns a list of 768 numbers. Those numbers work like coordinates: headlines that mean similar things land close together, even when they share no words. "Buyout talks with Acme" and "Acme nears sale to private equity" end up neighbours. That is what lets later steps search by meaning instead of keywords.

      # llm.py: Ollama serves the model on this machine POST localhost:11434/api/embed {"model": "nomic-embed-text", "input": ["search_document: Buyer agrees to acquire Target ..."]}

    3. Step 3

      Store them in a vector database

      The vectors go into Chroma, a vector database saved as files next to the code. Each entry keeps its vector, its text, and the title, source, link and publish time. Headlines stay for 30 days, because a deal plays out over weeks.

      This is the memory of the system. The language model never changes and never learns from these runs. What grows is the database, so today's note can draw on the story that first reported a deal two weeks ago, long after the three-day fetch window has moved on.

    4. Step 4

      Find the deals in the headlines

      The model reads the headlines ten at a time and returns, for each one that reports a deal, the buyer, the target and the value, in a JSON shape Ollama enforces. A model this small will sometimes invent a company, swap buyer and target, or make up a number, so the code doesn't take its word. Every name has to appear in the headline text, and so does the value, or the record is dropped.

      The records that survive are grouped, so five headlines about one takeover become one deal. If the headlines disagree about who is buying, the majority wins. The stage comes from the words in the newest clear headline ("agrees to acquire", "completes", "called off", "regulators"), because in testing the small model called almost every deal completed. Here is what was found this time, with the headlines each deal was read from:

    5. Step 5

      Search the database for each deal

      For each deal the code writes a plain-English question in the words a headline would use, embeds it the same way, and asks Chroma for the six closest headlines from the last 30 days. Closeness is cosine similarity: 1 means the same meaning, 0 means unrelated.

      A vector search always returns its nearest results, even when nothing is actually near, so anything below a similarity cutoff is thrown away. So is any hit that names neither company, because to an embedding model every merger headline sounds a little like every other one. These were this update's searches:

    6. Step 6

      Build one prompt per deal

      Each deal gets its own prompt with only its headlines, up to eight: the ones that revealed it plus whatever the search found, numbered the same way as the source list. A small local model stays on topic far better with a handful of relevant headlines than with a hundred mixed ones.

      The rules come first: use only these headlines, ignore any that are about other deals, cite by number, match the stage the code decided, and say so when nothing is new. A pinned deal with no coverage skips the model, and the page says there is nothing to report.

      # one of this update's prompts Deal: Buyer and Target. Headlines: [4] Reuters, Oct 5 09:12 AM ET: Buyer to acquire Target ...

    7. Step 7

      Write in a fixed shape, then check

      A Llama model running on Ollama answers each prompt at a low temperature, in a JSON shape Ollama enforces with a schema: a subheading of up to eight words, then two to four sentences, each with the numbers of the headlines that say it. The deal value, the stage and the article counts come from the code, so the model never retypes a number.

      A last call reads the deal summaries and writes the headline sentence at the top, then picks up to three other deals worth knowing, from headlines that aren't about the deals shown. Finally the code checks every source number against that deal's own headlines, removes any that point somewhere else, and drops a sentence that is left with no source. That catches invented citations, though it can't prove a source supports its sentence, so the sources are listed above to check.

      # final.py: Ollama holds the reply to this shape "format": {"type": "object", "properties": {"headline": {"type": "string"}, "sentences": [{"text": "...", "sources": [3, 7]}]}, "required": ["headline", "sentences"]}

    A GitHub Actions job runs the pipeline each weekday at about 7:15 AM Eastern. It installs Ollama on a fresh machine, downloads a small Llama model and the embedding model, runs the same Python, and commits the new deals file and the vector database to this site's repository, which redeploys the page. Models are set by environment variables, so the same code runs with a bigger model on a laptop. A deal can also be pinned to a watchlist file, so it stays on the page on days when nobody writes about it.