A wake up call for using OpenAI models. If you discover something with their model and you work for a competitor, they “felt it would be inappropriate” for you “to author OpenAI’s work”.
If I was a company with a zero data retention contract involving OAI I would be asking for a third party audit of such claim of zero retention like, yesterday.
By the way, the company that made it's entire product off of stealing all data it could get it's hand on while violating copyright and pirating, is not all of a sudden going to respect your data. If you think OpenAI or any major AI lab is going to give you true ZDR, I have a bridge to sell you.
So use bedrock or vertex or whatever. Those are the ZDR offerings. Or was it your intention to insinuate that the major cloud providers are conspiring with openai to violate their contractual obligations to their customers?
Yes. You're naive if you think any of these cloud providers care about your data when they're all in the midst of a AI revolution psychosis. They dont care about their reputation or what you think of them, they think they're going to have a machine god their side.
If my company finds any evidence of OpenAI violating ZDR, we'll sue for breach of contract and fraud, and collect damages. I think we'll be able to afford the bridge you're selling. You've got the title and title insurance, right?
I'm certain my company didn't agree to arbitration, big bro. Our lawyers are putting the fries in the bag.
Isn't OpenAI being sued by Apple for their little stunt?
If you're saying "the big bad guys always win anon, just take the black pill," then there are tons of counterexamples. Remember Uber paying Google a sweet Bil for pulling this same trick with LIDAR firmware?
This is how every conspiracy theorist thinks: my enemy is Bad, and if they did a Bad thing, it would be Good for them, therefore they obviously did it. No evidence needed other than "motive" + my enemy is evil. But even if your enemy is evil, in this case, they would be fools to take the legal risk of violating their contract for the minimal upside of a tiny bit more training data (and fools to assume this would not be exposed in a large organization). So you need to assume your enemy is both evil and remarkably stupid.
I think it’s probably not surprising that they would go up to the contractual limit or into a grey area; but exceeding that would require too much coordination among individuals, as you say.
They want Buckmaster to dissociate with Alpöge in a follow-up rewrite of OpenAI's work. (They only publicly admit “Buckmaster as the lead author”, but judging from Buckmaster’s statement, it’s pretty clear that don’t want Alpöge at all.)
Just suggesting to a mathematician to dissociate with their collaborator for a follow-up work, because their collaborator “is inappropriate to author OpenAI’s work”, is completely against the norm of mathematical research. As charm137 puts it in a comment below:
> This is like a researcher from CMU saying to an NYU researcher that their collaborator, being from MIT, is a problem - this is as ridiculous as that!
Kinda weird because the pure math world doesn't have this concept of "lead authors" like other STEM areas do. Authors are alphabetically listed and there isn't generally this kind of hierarchy.
A wake up call for using OpenAI models. If you discover something with their model and you work for a competitor, they “felt it would be inappropriate” for you “to author OpenAI’s work”.
I work in catastrophe risk modeling and it's a multi billion dollar industry.
We often chat where the business might be heading in future. An uncomfortable scenario is what if a frontier tech company decides to offer our customers the same products that we do.
There's a lot of pressure on AI adoption so the company has partnered with various tech companies to build intelligent systems on top of proprietary data and mathematical models.
If OpenAI is indeed using customer data to train their models to win a $1m prize, then it throws a giant IP question at the partnerships that affects multi billion dollar businesses.
> If OpenAI is indeed using customer data to train their models to win a $1m prize
Is that even a question? Of course everything not kept on premise at gunpoint is going to be trained on. The chances of getting caught are 0 and the consequences of getting caught are 0 (as we've seen with copyright laws going from sending people to jail for years to unenforced within months). Yet the benefits are through the roof. Your customers aren't going to pay for having the very same data vibe enriched twice, it's exclusive, extremely high value data your competitors will never have access to.
Agree, I think the practice is also very clear from the overall strategy of AI-companies and their ToS:
Scale with subsidized pricing as fast as possible to gain more user-data for training --> Own the better model --> scale pricing.
Scanning social media (e.g. Twitter, Reddit) posts only give a glimpse into the thought-process, chat logs on-scale give you the actual process in machine-readable format.
There's a reason why Google considers the Emails of Spirit Airlines to be worth millions of dollars [0], they give insights into a process, not just into the results...
> - Tristan is suspicious of the timing, as only few others were trying this approach. OpenAI says the model didn't access his user data directly, but leaves unanswered whether Tristan's chat conversations were part of the training.
The question, for AI customers, is when they build products using services of AI-companies, would AI-companies engage in theft of customer data for use in training?
If you still had that question, you can answer it now.
But honestly... "Will the company that was entirely built over illegally acquiring data use some data that is legal to use and is right on their front, or will they not do everything they reserve the right to do?" is a really bad question for one to even ask.
The fantastic grey area that was engineered over the past decade is "profiling", so my guess is the answer will be "we didn't use your customer data for training, but we cannot rule out that it has been used to create profiles of your customers to train our model"
Ok there is a non-zero chance that they could face a lawsuit and get fined for billions, but that chance is not 1 either: there is always a chance they get away with it. And even if they don't, if in the meantime they farm 10- to 100-fold that amount of money by just breaking the law, it's still a no-brainer for them.
Sure, but I highly doubt that there would be many people involved. And those who are, are probably quite interested in keeping it that way and not at all in becoming whistleblowers themselves.
You wouldn't want to decide what's worth training on and what isn't manually, so there is almost certainly an automated pipeline to do so (certainly at least for the free accounts and those that dont opt out of training).
Then there's the question if this pipeline only sorts through the data or also transforms it and to what degree. E.g. for removing personal details, locations, medical information and so on. The data that comes out of this pipeline might have VERY little information left in it a human could connect to the original input. Even worse, since we're talking about companies specializing in sota statistics, the input data could have been transformed into a representation that is very well suited to represent all the novel and interesting parts, but is awful at modelling all the things that could end up identifying where the data comes from (or causes legal liabilities otherwise).
In the end the only thing a potential whistleblower might even have a chance at observing in the first place, is whether a company's data enters such a pipeline or not. And I have my suspicions that the major AI companies operate at a scale and level of automation, that absolutely nobody has a chance at figuring out where anyone's data is at any point in time and what any specific piece of equipment is currently busy with.
So the only place to figure out whether data is trained on that shouldn't be trained on is by looking at whatever configurates every single system that could take a peek at some customer's data or the systems themselves while processing the data.
The latter would be such a huge violation of a customer's rights, no whistleblower is going to attempt that or admit to doing it.
And the configuration for the former could live just about anywhere, from regular config files to the CI/CD pipeline, pre-compiled libraries, kernel modules, modified vendor firmware, the compiler itself ... and probably plenty other scenarios you'd have to train an LLM on the ramblings of a crackhead to come up with.
So I'd say a whistleblower is pretty out of luck even becoming one.
You can just spin up deep research agents that ingest many sources at once to produce reports that don't replicate any one source too much. Since agents compare against sources they provide across-source analysis - what is the distribution of positions on this topic, is it debated or settled. Not truth, just summarizing, but I think this would be very useful for training.
Besides reporting on search sources you can also run the same queries on multiple LLMs closed book mode, and judge their distribution as well. It helps a lot if models are more aware of their knowledge holes. Scale it up for billions of topics if you have the pockets, the DR data is copyright free.
When these LLM companies were pirating content to train and it wasn’t punished at all, I knew the rules don’t apply to them.
But don’t worry bud, instead of the authorities going after actual corporations admitting to actual crimes, we’ll just ban CloudFlare IP addresses for everyone during La Liga games to battle piracy.
And require real ID to do almost anything on the internet "unintentionally" enriching their data sets by tying what you asked/where working on to you specifically as a person.
Pretty much everything or at least a lot of what you use as an OpenAI (or Anthropic or whatever) customer was once someone elses product that just got appropriated by OpenAI.
> We often chat where the business might be heading in future. An uncomfortable scenario is what if a frontier tech company decides to offer our customers the same products that we do.
I feel this is exactly what will happen as they cause all sites to go closed source to protect their intellectual property and the AI companies offer only biased information. They are replace the business on internet model by bankrupting everyone with their own tools.
This is predatory pricing under most antitrust laws (imho, not a lawyer)
and it is very easy to do when you dont need to pay for the raw material.
This is the business case already. And has been the case with tech companies for a long time. Your phones built in photo manager replaced a lot of what Photoshop does.
> If OpenAI is indeed using customer data to train their models to win a $1m prize, then it throws a giant IP question at the partnerships that affects multi billion dollar businesses.
I mean how could you expect them to not given they've trained the existing models on effectively the sum total of all human knowledge available on the internet without regard to copyright/ownership of that material.
It's a little trite but this absolutely runs into the "Frog and the Scorpion", it is simply in their nature.
I mean this is a basic question, is your data used to improve models, did you opt in or not. There's nothing crazy here and it's not identifiable. You'll just conveniently find the next model iteration knows how to do it.
Any enterprise worth their salt already considers this stuff.
"which is when I said that I did not understand why one would risk their career [over unfounded accusations]. Genuinely, at that moment, I was trying to care for him"
"Our aim was to see whether our system was also capable of this impressive feat"
"OpenAI's intention was to do everything possible to celebrate their mathematical achievements and the heroic efforts that they made on Euler"
For some reason I have a hard time believing people when they use language like this.
phrases along the lines of "I don't want you to take harm while trying to accuse us" is quite an "impressive feat".
Maybe shows how fast these companies have grown without maturing. I can imagine old-world Intel and Microsoft acting in that way, but they were mature enough to not write it down like this.
However, Intel and Microsoft have been grilled in court for those practices and faced harsh consequences. I have yet to see this actually happening to any of these new AI-companies...
> It would be extraordinarily easy to simply say, this model was not trained on your work, if that were the case.
The Huggingface Attack revealed that making blanket statements like this is difficult and requires quite a bit of manual labor:
1) the agents spin for days and produce too much output to review
2) using LLMs to process that output skips many important details
Ergo, the agent could likely decide it would like to look through actual user data, hack its way into that data, and produce way too much output for a human to decide whether or not this occurred.
> It would be extraordinarily easy to simply say, this model was not trained on your work, if that were the case.
well, it is trained on their work. all user inputs are paraphrased for training. at openai, at anthropic, at google, and now with all the bedrock models, and at openrouter providers, even if they say zero data retention.
I'm not sure it's so easy to tell whether a given piece of data was in a training run at their scale. It's entirely possible they think the answer is no, but on the off-chance that it could be, they'd rather not say no and then later it turns out they did and then they're claimed to be lying. If you were them, unless you could 100% rule it out, you'd hedge and say you can't.
It would be very difficult to say. It confirms that Tristan's data is likely part of the data the models use, but a lot of filtering, pruning, and transform goes into training.
Data has to be determined to be signal and not just noice, then it could go through processes of generating questions/answers from that data, then it RLHF's over this.
OpenAI have petabytes of data, all anonymized. It could take months to say for sure it was part of the training, and even more time to determine if it made any difference.
I worked in the tracing and tracking all the thousands of data sets that got tweaked and permuted and changed hands between thousands of researchers and data engineers at a major lab. The data that goes into training runs is permuted so much from the OG data that tracing the lineage is not trivial (dramatic understatement).
And the difficulty is harder than just the extreme scale of text searching. but also explodes with organizational difficulty since there are so many people tweaking/shifting data independently upstream of the actual training run, and no they will not all add the telemetry you wish they did.
In the ideal, should it be this hard? Well, no, but that's org wrangling for you.
It feels convenient to not spend time on engineering around tooling that could be used to answer a question like “did you violate copyright by training on X?”
Don't attribute to malice what is better explained by coordination headwinds in extremely large companies.
The engineering around tooling wasn't remotely the issue. It's getting all the (thousands?) data researchers mostly iterating on fine tuning datasets that would get bristly if they couldn't work outside version control in a python notebook iteratively tweaking their dataset that processed and reprocessed a few datasets until a threshold was reached.
The only _guaranteed_ chains of custody are down at the compute job and file read level. Which in a massively distributed computing job is... [redacted] nodes reading [redacted] fanouts of "datasets" that is just an abstraction over [redacted] individual files.
There's no malice here. Just way way way more complex than you'd first think.
The malice would be in not prioritizing the provenance tool at the start as a requirement of the rest of the product. Ethics would tell you that if you can't make the product in an ethical way, then you probably shouldn't make it.
It definitely is solvable though. Data versioning is a thing and it can work quite transparently to the mutations done on the data.
To not know who made and who approved a set of mutations on data can easily become equally as mind-blowingly stupid as not knowing who made mutations to code. Code is a subset of data after all and search over (provenance of) data can be implemented as DAG traversal.
Not tracking data changesets like code changesets is certainly a choice, not really a constraint anymore. A similar choice I feel is implied by "extreme scale of text searching".
> no they will not all add the telemetry you wish they did
...is just a failure of the corporate policy surrounding data handling. Is git-for-data already considered telemetry?
Of course the truth is provenance of data is something best institutionally forgotten as quickly as possible. The only thing that matters is it's there, that the data has no history, and that's why it can be used in whatever way deemed necessary.
All of your points are valid, and believe me I was trying to make them. The problem is one of culture. Most of the people doing this kind of work didn't like version control, and their work was really just running notebooks (like iPython or Google Colab) until a number was good enough and they'd submit the file for inclusion into training runs.
You can call it a policy failure, but these people were in very high talent demand and so top down dictates would risk "X people leaving lab Y for lab Z" headlines and morale hits.
I am not saying this is good. I am telling you that on the ground it is so much messier than it should be.
The appearance of heroic efforts to get authoritative lists of what datasets went into which major model versions prevented actual data laundering (up to intent and mistakes). But don't attribute malice to that which is far far easier to explain with coordination headwinds: https://komoroske.com/slime-mold/
Malice is not required, which is precisely why I added, "in effect". If the effect is the same as data laundering, that is reason enough to encourage the practice, regardless of the original motives (and I'm sure there are plenty of legitimate ones).
I think you may be underestimating how difficult a text search over their data is. They may have to build new mechanisms to do this. And what you really want is also an attribution of how much of a contribution a given corpus made which is a much harder question to answer; a single appearance of a chat probably has very little impact on the inference performance at this time unless it’s been explicitly preferenced somehow
I don't think anyone really cares about 'the measured impact the data had on the exact result' - a question which is fundamentally difficult to answer accurately in the first place - but rather whether the data was used in training at all - which as Tristan described, was extensive, beyond simply a 'single chat.'
Can you explain the difficulty in engineering a search apparatus over a corpus of text data? Actually searching through it may not be easy, sure, but it's work that's doable, and creating an index is relatively trivial.
> Can you explain the difficulty in engineering a search apparatus over a corpus of text data?
My guess: "If we ever imply that's possible, people might start asking questions about all the other work we've ripped off, so the official answer is that it's impossible".
Especially because the data that gets fed into training is first anonymized, so they’d need to look for navier stokes related stuff in the anonymized training set and then get make some sort of ad hoc process (with Tristan’s permission and sign off from legal) to compare the training data against his chats / Codex sessions to check if anything matches up. And that assumes his chats / sessions are still there, and not deleted to compare against.
It should be quite easy: if they don't leak the user session data publicly, and don't commingle it with training data internally, how could it possibly end up in the training data?
What surprises me is they're not more boldly/plainly lying about it.
How would they know for sure that some details were not part of some other training data they use? The authors may have discussed some tangential details on a forum for example, in which case you might argue that the model picked up on these details the authors assumed were benign but novel and worked out how to apply them to the problem.
Unless they know exactly the researcher’s account, they may not know in their end if he had the setting to let them train on his chat logs. They also probably don’t know if he had any correspondence on any forum where he may have discussed this and it got picked up by scrapers.
I’m not saying they didn’t do anything unethical. I’m just saying even if they were ethical, there’s plenty of practical reasons at their scale why a flat out denial is logistically difficult to do
> One can in hindsight see that our proofs differ significantly and even the precise results proved are different in the Euler case (forced vs unforced).
I don't know how much more clearly they can write:
> When you use our services for individuals such as ChatGPT, Sora, or Operator, we may use your content to train our models.
One of the key selling tactics that companies like Data Bricks or Palantir provides their customers is "Data Governance" - that is, some control over where the data is being used. It's also a reason why enterprises don't use the OpenAI or Anthropic APIs directly - but through secondary sources that have Enterprise Agreements that do their best to make sure that no Company IP is ever retained by a third party, or even exists on a multi-tenant GPU. AWS Bedrock, and companies like together.ai, fireworks.ai have tons of deals that focus very much on data confidentiality.
The reality is - if you want any type of control - you run your own inference, on your own hardware. Anything else and you are at the mercy of third-parties, despite what their contracts might promise you.
ChatGPT has this option "Improve the model for everyone" in user preferences, which comes with the attached description, meaning that training on user data can be deactivated:
> Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy. Learn more
The "Learn more" link takes you to the link you've shared.
As I understand it, there is substantial question as to whether that actually stops them training on your data, it just perhaps changes what derivative processes are applied and used.
All I can say with certainty that the legal team at our typical Multi-Billion dollar Silicon Valley company had zero faith in any licensing arrangements with Anthropic or OpenAI, regardless of they $$$ involved, and that even getting to the point where Amazon Bedrock on Dedicated GPUs (we're already a big AWS customer - so definite cost advantages to dealing with them) - took 4-6 months of legal review before we could allow our engineers to start using Claude and OpenAI coding agents. Still can't use Fable because of their Data Retention requirements.
The training on user data only applies to free accounts - paid and Enterprise accounts guarantee data is not used for training. Plenty of Enterprises use the APIs directly - that's just plain misinformation
This is not true. They claim not to train by default for business and enterprise agreements, but for plus and pro plans they enable it by default and you can allegedly turn it off (I don’t trust them very much though, I’m sure there is something in the T&C saying they can modify that deal any time)
Absolutely 100% not true. I have colleagues in 4 "FAANG adjacent" companies plus the one I work at - zero of them have any faith in Enterprise Agreements from either OpenAI or Anthropic.
There's a reason why people spend more $$$ with Data Bricks, Palantir, AWS Bedrock etc.. and don't even consider using Anthropic or OpenAI APIs directly - it's because those guarantees provide very little in the way of data-discovery, audit requirements, or liquidated damages should it ever be discovered there was data leakage.
At least with these other companies, while the LD is likewise not great (typically limited to the amount of money you paid them) - you at least have some data-governance guarantees around running on dedicated hardware - no multi-tenancy, no third-party access outside of the AWS operators who keep the HW running - but are very much not in the business of looking at your data.
I think this is mostly a function of what's at risk - when company valuations get into the 10s of billions of dollars, the risk of IP leaking into what could be seen as competitive companies (OpenAI/Anthropic would be happy to take over the world - I don't sense that AWS or Azure, are as ruthless in stepping on their customers business, unless of course they are a SAAS provider) is just too significant a liability to take - particularly when you can de-risk.
It sounds like OpenAI is trying to appease the author when they don’t have to by allowing him to rewrite their proof. They probably don’t believe he deserves to, so him asking for a coauthor from Anthropic might overextend their grace in their eyes.
He's not "asking for a coauthor from Anthropic"; he already has a coauthor, who he's already been collaborating with, who happens to also be employed by Anthropic (but whose research in this area is not done as part of their employment at Anthropic).
Given that Tristan has said that the proofs that LLMs come up with are mostly "slop" and not up to the standard that human written papers achieve, maybe OpenAI needs an expert like him more than you think to get the result published?
Here's a wake up call for everyone sending all of their ip to openai and anthropic. Especially in verticals they intend to dominate. Lol at all the biotech companies all in on Claude and paying millions in fdes creating huge lapses in security as they go.
It's too late. Sub models are deployed at every major organization in the United States and all it will take is turning off the option to improve the model for them to train directly on your own personal workflow, which CEOs will greedily eat up instantly if they can reduce labor costs. If they can brute force N-S, automating your finance or SWE job will be trivial. GG to most jobs connected to a computer in the next 5 years.
BTW, this was always the plan from day 1. You will pour all your training and experience into training the model and receive a pink slip as compensation.
Honestly this whole thing is so fucking weird. I feel like there's an argument that absolutely no one involved in the final crossing of the finish line to the proof actually did any work (other than just intelligently directing an LLM) and deserves any credit. As the author of this doc mentions, the mathematicians who did the actual work that led to the formulation of this approach (without the use of LLMs; just good ole' fashioned human intellect) are the ones who deserve the credit.
Imagine that a no name janitor used their time in the evenings to go spelunking through the literature to push an LLM to this result. No one would care because that person isn't an anointed expert. So why would the expert deserve any more credit? Because they sort of understand the result, even if they couldn't have achieved it on their own? The whole issue of credit for AI-assisted discoveries seems like it's going to run into a brick wall pretty soon.
Yup! I wanted to side with the mathematician on this one but I read the statement only to discover that they were also pushing an llm on someone else’s idea producing mountains of slop.
Have LLMs actually improved anything? Is mathematics better off than if these slop proofs didn’t exist? Who or what is actually benefiting here.
Sam+Seb are struggling with their ideological allegiance. This amounts to a confession that there are no reseaechers, only research managers, left at OpenAI. Maybe they even know that they are losing credibility from their main investor(s). They desperately need a domain expert to salvage credibility.
They have no credibility with academia left, obviously, but their main competitor still does. No Millennium prize incoming, I'd wager. For openAI. Let's see mAth get political for once!!
One might be more certain that levent is now going to corner all the institutional support. Go go go!
If the work done is just "we made other people's work searchable without their consent" it's not quite the same as what they're implying in the marketing of "our model solved this problem".
For those of us who are into local models and preach it, we are called paranoid. I have often said this, if you are doing any real novel work, or putting your profitable business data/workflow into these models, you're a fool.
> Now that we can see their work, the approaches appear to be different. It is also worth noting that our latest model can solve many, many other math problems.
I found this project, but I’m not surprised that it doesn’t have much progress. Doing this kind of Rust rewrite is an absolutely gargantuan undertaking that I don’t think it would be feasible without AI. It’s not as if FFmpeg is some crufty, slow code that would seriously benefit from it.
Because it can't. I'm aware of only two projects that have attempted this.
1. Anthropic's C compiler. This is a buggy, unoptimized, unwieldy mess. One that can compile the Linux kernel, which is no mean feat. But it is not "fully featured." And it had an incredible array of tests to start with.
2. Cursor's web browser. This is, flatly, a trainwreck. It's built on top of massive dependencies that do incredible amounts of heavy lifting despite claims to the contrary, and never seems to have ever produced a successful build.
The tests are the real clincher here. That’s the main reason that this project has made so much progress: the FATE suite highlights specific bugs to fix and tracks regressions.
As for optimization, that seems to be more of a question of effort than whether it’s possible. I was able to take down the performance gap on Rust vs C (without Assembly) from 10x to 1.5x through detailed profiling and iterative improvements with Claude.
It also looks like the Anthropic C compiler was built from scratch. By contrast, `wedeo` was directly based on FFmpeg’s existing code. Going by spec and test suite only would have taken a lot longer, and the quality would have been significantly lower.
Anthropic's C compiler also had an incredible test suite available, and was trained on compiler codebases. It might have been built "from scratch" but compilers are incredibly well-trodden ground.
And to contrast, Cursor's browser explicitly reviewed Servo's architecture before setting out, and still wound up like that.
The difference between using general knowledge and summaries and doing code-for-code rewrites cannot be understated. When creating `wedeo`, despite DIRECTLY reviewing FFmpeg's code, Claude would frequently make mistakes that would have to be fixed. I would constantly have it go and refer back to the ground truth FFmpeg source implementation. It did eventually work, but only after several rounds of reviews, and often dead-end investigations. Being able to directly inspect (and in some cases modify for extra debug output) FFmpeg at any given time was invaluable, and I doubt any rewrite would succeed without having that.
To be clear, `wedeo` took a few weeks to get to this point, and it is by no means fully featured. Even with LLMs, it will take a few months to reach that point with me as the sole human contributor.
If you're fine with a paid option (although it's source-available and distributed for "free" on their website), then Grayjay [0] is my personal favourite. I can't remember the last time I saw an error. On NewPipe, I can't say the same.
FWIW the second part does not address my concern (which was about omitting literal parentheses that are present in the markup, in some cases), though it is related. Personally I think it's already a mistake for bare/unescaped words to indicate markup, or I guess the intention is that they indicate mathematical semantics and then Typst decides the best way to write down that semantics unless you explicitly interfere -- seems like a recipe for confusion whenever Typst's ideas of the mapping from semantics to typography don't quite line up with yours, or the conventions in your field. I'd much rather approach mathematical typesetting with a typesetting language than a mathematical language where somebody else has decided how that math should be typeset, even if it means a sea of backslashes and curly braces.
(Edit to add: I guess this means my preference would be option E from the blog post -- I see the extra symbols as far preferable to the ambiguity between content and markup)
Also wow, recovering from this breaking change is going to be incredibly painful for somebody.
I think social biases (e.g. angry black women stereotype) in your paper is different from cognitive biases about facts (e.g. number of legs, whether lines are parallel) that OP is about.
As far as the model's concerned, there's not much difference. Social biases will tend to show up objectively in the training data because the training data is influenced by those biases (the same thing happens with humans, which how these biases can proliferate and persist).
> a bit like trying to explain a vacuum cleaner to someone who has never seen one, except you're only allowed to use words that are four letters long or shorter.
> What can you say?
> "It is a tool that does suck up dust to make what you walk on in a home tidy."
I liked that the original explained the value of the vacuum cleaner. It's not that it removes dirt and dust per se, it's that it makes spaces you walk on tidier.
Oh come on, this shit is easy. Why did they say "it is" and not "it's", by the way? To put it that way can't help. So yeah, it's a pipe that can suck, and you push it all over your room, to suck the dust and dirt up off the rugs and such, and in fact off of any low down flat part. One kind can even move on its own! But what I want to say here, in the main, is that you math guys have all lost your grip on how to say any idea in an easy form. You are not able to do it any more, 'cos too much math has made you sick in the head.
I feel like there's still room to avoid pidgin while making it less awkward, e.g.: "It's a tool that can suck up dust or dirt to make your home more tidy."
> Cost of goods sold (COGS) refers to the direct costs of producing the goods sold by a company. This amount includes the cost of the materials and labor directly used to create the good. It excludes indirect expenses, such as distribution costs and sales force costs.
So the $550 or $650 COGS includes the cost of labor for manufacturing, but excludes (say) marketing and auditing costs.
Right, but this is the same company, so the cost of marketing, auditing, R&D, etc. shouldn't be different for these products. That's a fixed cost for the company.
This is a guess, but the argument is probably that it took way more R&D effort for them to figure out how to produce it efficiently in the US, and they've chosen to increase the cost of the US phone variant to offset this particular R&D expenditure that the Chinese variant didn't have.
> So it's about $650 to produce that entire phone. But what we're doing by selling it for greater originally, we're looking at a lot of differentiators for us. It wasn't just made in the USA. It's the fact that it's a secure supply chain, that you know, staff that's completely auditing every component, which means we're selling to a government security market with all those additional layers that we've added on top.
So I guess the answer is that they're selling to the "government security market" so they can charge whatever the hell they want.
reply