To me it's a good replacement for search engines. I ask it things like 'If redshifting destroys energy ala Noether, than how can we say that time is reversible or that entropy will find an equilibrium?' and it will not only explain, but make nice interactive diagram/toys to help. A regular search engine would have taken me hours to find an answer.
However, if I want it to DO something then Gemini is in absolute last place. I don't trust it for anything more than renaming files that I don't care about very much or extracting data (though it's too expensive for data extraction at scale).
With every VR release I check the resolution per eye to see if I can use it as a work desktop replacement. IMO the current Meta Quest 3 is not there, and this roughly matches it. I am not sure what the resolution needs to be to render good text...but I'm hopeful we get maybe 50% more pixels at some point in some version. Unfortunately, i think the screen stack at that level is well outside mobile usecases so there is kind of a chicken and egg. Manufacturers don't want to work on advanced VR display tech without demand, and it can't be used for work purposes until display tech improves.
We had a two human PR requirement until recently we dropped it. It was slowing us down too much now the human developer creating the future is obviously writing it all with AI so they need to check it then depending on the feature and it’s use it requires a PR but it’s not universal and we’ve stepped up our automated test Tan X what it used to be it’s been so far fewer bugs better delivery
If you imagine someone speaking that comment out loud, but speaking as if they were giving a keynote at a Meta or Apple dev con, it becomes much easier to read. The commas and dramatic ellipses just fell into place as I read. Like the matrix, but instead of green kanji raining down, it’s readability-increasing punctuation.
lol
Certainly not always. There's a hedonic adjustment which happens however, where some tasks go very smoothly without much specification and a lot of "you know what I mean" to the LLM, while others then require you to get painfully specific after it badly misinterprets your intent.
Or maybe you can just get too spoiled with it grokking your intent, then become so vague that your vague ideas are actually just bad ideas. Certainly has happened to me.
After almost 4 years with LLMs, if my prompt is too vague that I don’t know how to ask precisely, I use this prompt “I have this idea… {description here} How would ideal prompt look like to make idea realize?” And in second round I polish prompt myself. It usually works.
Apple's search is so comically bad. For example, I can click on my applications, and I have an app called Chess. I can often type letters C-H-E-S-S and have it not show the Chess app. This is inside the applications folder!
Many people have complained before. Honestly I ONLY search for app names in spotlight search or application window, it should be trivial to make it actually work. Maybe i just need to vibecode some alternative search bar :(
It’s so much worse than that. This happens to me on a daily basis with both MacOS and iOS. Say I’m searching for Chess, I’ll get to the C-H and it’ll show the Chess app, but enter that next -E and it’s totally gone. It feels like the system is designed like “Well, I showed them Chess when they entered CH, but they kept typing, so they must not actually want Chess.” This ends up breaking for me like 95% of the time, and the only way around it is to search like “C”…wait for results…”H”…wait for results…etc. It’s maddening.
This, and then, if you notice, and backspace to remove the 'e', it doesn't put chess back into the search results, since it figures that everything that matches 'che' also matches 'ch', and so doesn't update it. Sometimes, if you add the 'e' again, it _will_ find chess, sometimes you need to type the 's', or even the final 's' before it gives you chess back again.
And sometimes, it finds what you want, but before you can select it, it decides that it has found something better, and places that in the place where you wanted to select, causing you to select the wrong thing. And then it thinks that because you selected this, it must be what you want, and starts to push it to the top…
Weird, the same stupid crap happens with Windows. I’ve had to learn the exact N character search to find the things I use often; any more or less and the result with that name is buried. It’s like Microsoft and Apple are just pip installing same broken search package.
I switched to Everything for search a while ago and haven’t looked back. For the start menu I replace it with Open Shell, which seems to find apps much more reliably.
Always turn off searching the internet as much as possible on both. Helps a ton. I think I set both windows and iOS to only search local applications and settings
Firefox's "awesome bar" behaves the same way. I guess it's either inherent to the problem somehow, or else everyone is sharing a library (unlikely), or else everyone just makes the same mistakes.
It's worse than this! You Type C-H-E then you go to press enter, and it changes the results as you're pressing it, opening the wrong thing! I've turned off all the search targets, except for applications and folders.
Yes. "UI Speed" should be a configurable thing, so that if the UI changes as you're choosing an item, it acts instead on _what was there_ at a brief moment in the past, e.g. 20-50ms ago.
Don't worry, airdrop does it too! I've accidentally sent airdrop requests to strangers from their profile picture moving under my finger as I'm approaching to touch someone else's. The single most idiotic design I've ever seen in UI. I don't use airdrop for anything sensitive anymore, obviously.
I really wish apps stopped doing that "search on type" thing in search boxes. You either need to have only a couple miliseconds of response time, or you need to confirm that you want the search process to start. Or maybe at least have some feedback showing "more content is loading, watch out, don't click".
The amount of misclicks daily because of crappy slow-but-should-be-fast search probably offsets by 10x the amount of time required to hit "enter" for the search to start.
For me search has been that way since the first MacOS 26.0 upgrade and apparently never fixed.
But speaking of quality overall (outside of the search glitch) - I use dual 5k LG monitors with my work MacBook and can get the MacOS Display Manager process to crash if I just enable variable rate refresh and put the computer to sleep, then wake it up. Apparently a known issue for a few years, still not fixed.
It's probably asynchronous reranking doing something weird. My guess is that the underlying ranking algorithms in these desktop search bars are very ropey, and they can't debug or diagnose because privacy reasons means the developers don't get to see what results were generated for different query/index combos. Also search is probably just not staffed most of the time.
So many platforms with search have a variation of this problem now.
The other issue with Mac search is that it never tells you when it’s done, so you can’t tell if it searched and failed to find something or is still searching.
'calc' used to bring up calculator, but now it becomes some widget to calculate distance inside the spotlight box. The worst thing is you can't even backspace out of it.
I'd take the widget! For me typing "calc" finds the "com.apple.calculator" folder. If I continue to "calcu" etc, it suggests "com.apple.ncplugin.calculator" instead. Brilliant intelligence right there.
For me the big one is whenever I type "sys" to get system settings, it pops up for a half second and then for some reason the top result is like "reset some system setting" and then I invariably get flashbanged because it turns off dark mode and there's no confirmation dialogue lmao. Maddening.
I've deactivated every single type of data in Spotlight except Applications. It goes out of its way to still suggest files like C system files or header files, like "system_library.c". It didn't wait for AI to turn against us.
I was trying to see what happens here. It does give the Sound settings panel as the first hit. But when I then press Return, it opens System Settings, but does not actually go to the Sound settings, but General.
Also, search inside the Settings app has been broken for many releases now. I always wonder if anyone inside Apple is actually using macOS or iOS.
I had this problem on Spotlight, I solved it once and never happened to me again, there is a setting to reindex all the files, which solves the issue. On iOS however I still have this problem.
I wish all this shit would just stop trying to search as I’m fucking typing. It’s like that person we all know that can’t hold their tongue for five fucking seconds and let you finish a sentence before the blurt something out.
Ah, but hitting enter now takes you to the top result! Therefore it has to search as you type, so there's something to activate as soon as you hit enter! /s
...I've been seeing this search style in more and more places recently, and I hate it. Let me type my search, then hit enter, then you can go find what I asked for. Don't slow everything down by searching on every character I type.
It’s called incremental search and it’s useful, but only when there’s some debouncing involved and the search results is presented at once or the items are sorted by the order they are found.
I can't search for Apples own Notes app. I've tried the name in my local language, I've tried the original english name of the app. Nothing. It will show me OneNote which isn't even installed anymore, but not the Notes app. Very strange.
- Or is the goal to disturb the user and make something impractical like in Google, so that people click on ads? Software has turned into a curse. The machines have risen.
I always assumed that was intentional to save you arriving through options. I find that behaviour quite useful as I can hit anything beginning with CHE just by typing more letters.
>> the only way around it is to search like “C”…wait for results…”H”…wait for results…
Some people don’t like that. I don’t know why you like having to wait but if they showed Chess for CHE you could still find other apps if you kept typing.
Maybe there are also people who like the lottery of results changing the instant you click on the app you want because it adds excitement to their lives.
FWIW, I find if I’m pushing the disk limit, Spotlight search becomes unusable. So I free up ~100GB (easy to do with iOS/Android dev artifacts, emulators, etc.) and then force a Spotlight re-index from the command line:
sudo mdutil -E /
You can check indexing status with:
mdutil -s /
After that it’s usually fine again until I start pushing the disk limit.
Every sibling comment's issues I've experienced. But a reindex after freeing enough disk space seems to be a semi-stable solution for me. If only I could write a tool to remind me to free disk space before it screws up spotlight. It would be nice if MacOS put the index on its own partition.
Spotlight worked best in OS X 10.4. It was near instantaneous on what is now considered ancient hardware. It's been steadily downhill from there.
Mostly because people still expect to use it to find "Recipe for Egg Salad.docx" without searching the entire Internet or executing some bizarre natural language query.
A common one I hit is "blu" showing bluetooth settings, but "blue" does not. Very frustrating to see the thing you are hoping for popup, but then go away as you type more if its name.
Android keyboard does that too. When I start typing my email address, for example, the keyboard app suggests the correct email after a few characters, but if I keep typing it disappears from the suggestions.
I have to send a daily email that always has the subject starting with "Work Diary". If I type "W", it doesn't come up. If I type "Wo" it comes up as the 3rd auto-complete suggestion. If I type any more characters, I have to type the entire two words. This is on Android, but I have a feeling it's no better on iOS.
I'd love to know the algorithm they're using for this. Fuzzy search seems like it was a mostly solved problem 20 years ago, and it's definitely worse than it was 10 years ago.
It would be interesting to learn what over engineered solution to fuzzy string matching is being used here to exhibit this behavior, so we never make the same mistake.
Seems like someone got updated KPIs and decided to try the most ham-fisted enshittification move to raise registrations. Oh, it is now possible again to see about five top reviews per item. Seems like loss of traffic was rough.
Just tried this on iOS 27 and both blu and blue showed me bluetooth settings. I saw someone in a different thread of this discussion mention resetting dictionaries helping with autocorrect, I wonder if there is something that would apply to this situation which helps.
At some point I got the damn Chess app every time I wanted to switch to Chrome. You know, typing cmd-space, ch and enter. I've learned I can't remove Chess from my computer. After numerous incantations of various spells and reboots and updates I think it started working normally again.
Breaking interfaces, Linus Torvalds had something to say about it that Apple should heed.
I struggle daily with Finder's search. It cannot find files in my Downloads folder even if I type the exact name. Do I need to rebuild some index or some other setting?
It couldn't hurt. I don't have problems like that.
mdutil -E -a
That command will erase and rebuild Spotlight index on all volumes.
Since Tahoe, Spotlight has multiple modes. You might have better luck forcing it into folder & files mode. That's command-2. So command-space, spotlight UI appears, keep holding command key, press 2. There's also 1 for applications only, 3 for "actions" and 4 for clipboard history.
The annoying thing with Finder is when you're in a folder, you start to search, and then it defaults to "This Mac" and not the god damn folder you're in.
That's because you have that set in your Finder Settings.
With the Finder open, go to Settings > Advanced, and change the drop-down at the bottom to "Search the Current Folder". This is something I do on the first run of every new Mac I set up.
On my incredibly midrange PC that has a midrange SSD in it, the Linux `find` and `grep` commands usually can find stuff much quicker than any 'accelerated' and 'cached' desktop search, even though the only caching they usually get is when the files are in the page cache because I have interacted with them. Even when that's not the case, it's faster AND more accurate that GNOME search or whatever search is in Windows Explorer.
Windows is the same - searching the fs is faster and more accurate than using these 'quick search' tools, which are borderline broken.
Search feels like it used to be much better in finding a balance between showing app vs files vs web - feels really broken somehow and I wonder if there’s any setting I can change.
The good part is you have the source so you can just fork it, fix the code to prioritize results properly, compile and install it to fix Ubuntu, then see if you can get your patch upstreamed. /s
For me it's gotten way better in iOS/macOS 27. It's also way faster. They claim to have rewritten the spotlight index to handle the new Siri integration.
The new Siri UX needs improvement and to match the UX of talking to chatGPT. At least in settings have an option to keep the Siri window always on top (just have a prominent X out to go to the screen below) and always listening ready to continue your conversation like the experience is with chatGPT.
Been running the beta since June -- on the latest beta now and 8 out of ten times to continue the conversation with Siri (most notably in the car with my phone docked) I have to constantly say "Hey Siri."
One thing it's current UX has over chatGPT voice tech as it shows you the text and provides visuals while chatGPT's UI shows no text and or visuals.
I did just that. Took me half a day. I haven't used Spotlight since. I did try alternative launchers, but they are all bloated. I just need an app launcher and simple arithmetic - that's all it does.
Check out TinyStart. It’s 5€, and the best 5€ I’ve spent in ages. It’s super fast, and handles basic math. Even lets you put the results on the pasteboard!
Spotlight works for me across files and apps on Mac, iOS, visionOS, pretty well - it has failed occasionally but usually it's because my disk is full. I haven't really found anything better (Alfred is good but is just a wrapper around Spotlight)
I quit Apple Mail (due to its poor handling of Gmail accounts) and started using Mimestream years ago. Truly excellent mail client. Actively and thoughtfully developed.
They did fix it (macOS 27, my experience). I have like 600,000 mail messages - there was a period where Mail search was awful but these days it's fast and accurate. It really comes down to whether you need to rebuild your index.
I’ll have it not show Xcode. It’s apples main developer ide. Most of the time it works. Then it gets into some weird loop. Only shows the Xcode website. Very helpful.
Raycast2 has its own index and it's way more reliable and fast.
I don't understand how Apple does not care about many of the core services of their OS. Searching in Mail is a coinflip
Ever since they re-did Spotlight I've had issues. When I type something and press Return it should activate. I shouldn't have to wait for it to populate the results and then press Return a second time. Moreover they've broken the ability for me to use Cmd to locate files quickly. I have to awkwardly Cmd-double-click the result with my cursor because trying to Cmd-Return triggers Ask Siri!
Turn off all the web search and other domains that the index uses. Apple took a really brain dead decision to include web search in the results, and they are the worst. It is completely useless, and even though the recall has expanded, the precision is abysmal.
Alfred is wonderful for launching applications and custom actions, but it’s actually worse than Spotlight search for searching for files and folders. It finds a lot of results but inexplicably it likes to rank partial matches ahead of exact matches.
There was surprisingly little information on how they actually do this, but we run our business with a large number of, what we call "AI employees" in addition to regular employees, and they act in interesting ways. We've been building out orchestration tools to handle this, and we'll probably do a write-up or blog post on it soon. May even open source some of it.
The preview is that the problem with most agents (and this includes frameworks like Grokbot and Openclaw and Hermes) is that for many of them, they're black boxes. They say they learn or improve, but it's a black box in what they do. Getting agents to reliably do things is hard, and getting agents to build out software tools to help themselves improve and do better over time is also hard.
Our approach at a high level is pretty simple: every single AI employee is a standalone GitHub repo that shares some characteristics, but we direct them to build as much software as possible to make their goal as easy and reliable to manage as possible. Then we have a shared communication layer for bots across the company to interact with humans and AI. We have decided to organize these like departments similar to the way you might hire out humans. I'm not 100% sure if that's the best approach, but I will say it's been easier for people to understand because they're more mentally easily able to traverse the bot org chart if it somewhat reflects a traditional business org chart.
Each of these AI employees has specific sets of goals and KPIs, instructions that they manage the business with manager bots. We have layers of management, which we actually have found helpful. We also run different bots with different models and harnesses, and some using different models and harnesses to check the work before anything can get done, along with lots and lots of testing.
Every single time, actions have a massive amount of tests based off of previous failures to prevent failures in the future. Sorry for rambling. I do think this is a very interesting space. I didn't see anything interesting in Pion that was public on this website, but I do anticipate that more companies will be "AI and software first," as in the substrate of the company is basically a software application powered by autonomous agents, with humans as a fallback.
The whole premise of AI is to replace the human in the loop. People who go out of their way to set up this kind of stuff rather than just hiring a person dont even think about just hiring a person.
In the end, most software is a necessary evil. It solves a problem that shouldn't be there. In the end it's no different than healthcare or prisons. We don't need as much as possible. We need as little as possible. This automation isn't actually helping us.
That may or may not be true, but the calculus has changed with agents. A process that was better manual for a human org may not be better for an agent org.
I think what is interesting here is that the industry is in this experimentation flux. Some people will choose to automate and write the software, others will not. And aggregate over the industry and over time we will learn.
So just saying "sometimes it works and sometimes it doesn't" isn't really adding value, compared to the people actually experimenting and sharing the results.
> We have decided to organize these like departments similar to the way you might hire out humans. I'm not 100% sure if that's the best approach, but I will say it's been easier for people to understand because they're more mentally easily able to traverse the bot org chart if it somewhat reflects a traditional business org chart
I think this way of thinking is going to be very important to drive adoption of AI systems because the human analogies benefit from the pre-existing domain knowledge and expectations of people.
I like using the exam analogy for evals as a qualifier for work for your "AI hires" so you can trust them to work on a specific domain.
I'd be quite curious to see what your approach to evals/testing/tracing and agent/system mutation is.
This is fascinating. What is the split between human and AI labor? What kinds of tasks are the bots doing? How much autonomy do they have to make decisions (i.e. spending money, issuing refunds, touch cloud infra, etc)? I'd love to learn more about how you do this.
I hope to post something within a few weeks. But our philosophy is generally if it can be done deterministically with software (eg run payroll on autopilot via an api to gusto) then do that. If it can be done by an ai agent, do it with that but add as much software as possible to make it reliable at that thing. And the ai agents are all tuned to escalate to humans as needed. Then there is also just things that are completely human.
A typical example of something that is AI vs human is the AI most commonly operates like “managers”. For example, reviewing transcripts of every demo call, compiling results, figuring out insights, learnings that need to update our company docs, feedback to humans (who run the demos).
We are big fans of having AI agents “own” koi’s because now anytime we say “we really should be doing this” we try to set it up on the spot.
The “downside” here is that I do occasionally get busy, and if I’m the only one who can approve or unstick one of these bots, it just keeps harassing me until it gets done. This is a sign generally that I need to hire someone to own a set of bots.
At $DAYJOB we do something very similar shape-wise. I wonder - it sounds like you have a dedicated agent comms plane? In our case we found that the easiest and most straightforward was to just use our default company chat app directly for this. Because most of the context that the agent workers need to do work is there, but also, it’s just much easier for teams to conceptualise an agent colleague if it just hangs out in their channels.
We do have a control panel, but it’s in effect a server that has things in a database. All chats, tasks, assignments are all there. This felt easier to debug and manage versus putting it all in slack, although we considered it. I don’t think anything we are doing is particularly “fancy”, but basically one agent sends a message to another one. It saves the message in the db, then adds the message to that other bot same as any other user chat. We have an internal website where anyone in the company can see the bits they are authorized to see and can see all chats. These AI workers are single threaded, but we see that as more of a feature than a bug (minimize complexity). They all work their way through a shared task list, which is just another table in our remote server sqllite.
What harnesses do you use? Ours is basically Claude in a box. There’s some complexity because of that, but the advantage is that it’s very flexible and people who have a bunch of Claude-shaped skills can just basically give those to an agent.
I’m thinking to take a deeper look at Pi. I’m really liking that project.
how do you decide to start a new agent, and when to kill? also... do you use some works-tealing style task board, or otherwise how would the agents get new tasks.
It can't be bargained with. It can't be reasoned with. It doesn't feel pity, or remorse, or fear. And it absolutely will not stop... ever, until you go ahead and come in on Saturday.
What type of roles "AI employees" play, can it be any position in your company or you limit it to something specific?
What is your goal, are you trying to find out if fully autonomous bots are more efficiently help to deliver projects than when people drive them or is it something else?
Can you mention some specific tasks/KPIs that you've found bots can do a large amount of work autonomously on? By large I mean something that would take Claude Code or Codex at least a few hours.
What worries me most are questions like this: these autonomous AI employees or automations - it doesn't change the essence - consume a lot of LLM tokens. Please tell me, how do you pay for this? Do you use, for example, the Anthropic or OpenAI API directly, or do you connect, for example, a Codex subscription?
Our control panel is on a server but individual agents actually are run on anyone’s machine. This allows us to use the native harnesses including subscriptions. Yes it uses a lot more tokens, but i have have Claude $200, OpenAI $200, and SuperGrok Heavy $300 (or whatever it is called) that includes Cursor Ultra. My machine runs most of them, but some other team members have agents running on their machines using 1-2 $200/mo subs.
I would say this setup probably costs us about $1,000/month total. (I’m excluding traditional engineering use of LLMs from this number. This is the cost of all the “AI employees”.
I've been spending all my time recently thinking about LLM costs. From that, I am currently of the mind that the most interesting question is if they can get it to work in the first place. After that, like with "performance" in the past, I believe there are ways to optimize things. We see the code harnesses doing this out of necessity, for example.
There is a point where those two lines do cross but with them effectively subsidising token cost by burning debt, it's further away than it will be at some point.
That said I use local only models purely because I don't want to use remote models, never having to think about token costs is worth it and no one is training anything on my data either.
Don’t you end up with useless slop? That was my experience with these kinds of systems. In theory they should work but in practice unless hand holded they just produce slop and waste.
Google display ads are completely unusable for me (not app ads). So much flagrant fraud. Like even the most simple of AI models could detect the fraud. Very disappointing.
I have a different perspective. Most of the world is software. The US Constitution is software. Companies are software. With AI, we can build so much more software, its incredible! How these systems are architected are complicated. Sure, if you really loved coding at an extreme low level, that may go away. But I think there is 10x more software engineering going on my company (and in my personal life). Things change, but things are more fun than ever.
It is correct to compare to average drivers because that is ultimately what they are going to replace. AVs are coming for the whole market, and if they are significantly safer than the average driver the average driver shouldn’t drive. I intend to stop driving when possible.
reply